Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 12961 - 12990 of 63010

Full-Text Articles in Computer Sciences

Uc-238 Sp4-Dogem, Chloe Chung, Zane Atkinson, Melissa Iniestra, Deo Intshakal A Nzeng, Joshua K. Willson Dec 2022

Uc-238 Sp4-Dogem, Chloe Chung, Zane Atkinson, Melissa Iniestra, Deo Intshakal A Nzeng, Joshua K. Willson

C-Day Computing Showcase

The DogEm project's overarching objective is to produce a functional cross-platform mobile app that reliably sends a specific contact a series of calls, emails, and texts until the contact responds. To accomplish this goal, we have compiled preliminary research, produced a series of prototypes, and started the development process. Our preliminary research consists of research pertaining to our tech stack, our user base, our app's requirements, possibilities for messaging and calling features, UX/UI research, reading through React Native documentation, and constraints on the messaging and calling features on iOS vs. Android. Each member produced a DogEm app prototype to familiarize …


Uc-263 It4983 Cybersecurity Capstone, Jordan White, Stephen C. Woodman, Jenny Owens, Hector Gomez, Aaron Scott Dec 2022

Uc-263 It4983 Cybersecurity Capstone, Jordan White, Stephen C. Woodman, Jenny Owens, Hector Gomez, Aaron Scott

C-Day Computing Showcase

For the Cybersecurity capstone project, our team was given a webserver and website for the company Akwaaba. We were tasked to fix all vulnerabilities, keep the server up to date, and help maintain the site with uptime being the priority. After the first two milestones were complete, we were to attempt to hack another team’s server while they attempted to do the same to us. When the attack phase began on Wednesday 10/26, our team discovered the other team had not changed any of their original passwords, so we took control within thirty minutes and took it down soon after. …


Uc-268 Buggy - Price Scraper Application, David W. Fitzgerald, John Blake, Anderson F. Guse, Jessica Zardoya, Paul Allen Dec 2022

Uc-268 Buggy - Price Scraper Application, David W. Fitzgerald, John Blake, Anderson F. Guse, Jessica Zardoya, Paul Allen

C-Day Computing Showcase

Buggy is a desktop price scraper application that finds the best deals for users across several major retailers. Retailers include Amazon, Costco, Target, Walmart, and eBay. It simplifies bargain hunting by condensing product information from multiple websites into one, convenient place. Users can save products for future access, track search history, and compare search results with tabs for navigation. When they are ready to buy, they can simply click on the buy button and are directed to the vendor's page to check out.


Uc-275 Network Simulation Software Analysis Of Alternatives, Justin Mccannon, Nidhi Marsonia, Michael Mcinnis, Tiffany Nguyen, Azm Uddin Dec 2022

Uc-275 Network Simulation Software Analysis Of Alternatives, Justin Mccannon, Nidhi Marsonia, Michael Mcinnis, Tiffany Nguyen, Azm Uddin

C-Day Computing Showcase

Data networking is a complex field. Designing networks is a complex task. Simulators have been developed to create hypothetical designs with configuration settings for evaluating architectures and settings. At least 2 simulators are freely available to IT professionals and students. This project involves researching the landscape of free network design simulators to determine how many there are, then downloading, testing, evaluating, and documenting the features of each by designing, on each, a network with multiple routers, switches, and host devices using IPv4 and IPv6, defining features appropriate for classroom use in a university, and finally determining which solution fits the …


Ur-280 Speech Recognition, Rachnicha Rojjhanarittikorn Dec 2022

Ur-280 Speech Recognition, Rachnicha Rojjhanarittikorn

C-Day Computing Showcase

This project makes use of speech recognition in Python libraries from Google, such as Google Speech-to-Text. The program first installs the dependencies that are required to run the program. The program then asks the user for their input. The user has the option of either exit the program, having the text read back to them, or responding to the conversation. If the user decides to have the text read back to them, the program will use the Google Text-to-Speech library to convert the text to speech. The purpose of this project is to understand how to implement speech recognition in …


Ur-285 Operation Enduring Freedom: Improving Mission Effectiveness By Identifying Trends In Successful Terrorism, Dalton A. Shaver Dec 2022

Ur-285 Operation Enduring Freedom: Improving Mission Effectiveness By Identifying Trends In Successful Terrorism, Dalton A. Shaver

C-Day Computing Showcase

This research examines how the characteristics of terrorist attacks predict the chance of an attack succeeding, where an attack is defined as successful if the intended attack type is carried out. Data was analyzed across three geographical missions within Operation Enduring Freedom: Trans-Sahara, Horn of Africa, and the Philippines. Using predicted probabilities of success obtained from logistic regression models, the medians were plotted to compare the characteristics of terrorist attacks across missions. By determining the specific features of attacks that produce the highest probabilities of success, the effectiveness of Operation Enduring Freedom can be improved by focusing counter-terrorism training and …


Ur-302 Using Quantum Computing To Determine The Optimal Path On Cascading Graphs, Michael B. Swann, Ethan K. Hunt Dec 2022

Ur-302 Using Quantum Computing To Determine The Optimal Path On Cascading Graphs, Michael B. Swann, Ethan K. Hunt

C-Day Computing Showcase

Quantum computing has completely changed the computing paradigm. These special computers leverage the unique properties of quantum mechanics to solve problems that a classical computer cannot solve in polynomial time. Quantum mechanics such as superposition and entanglement are used to boost computational power exponentially in many problems . Many traditionally NP-complete problems, such as breaking the encryption of public-private key systems, are solvable with quantum computing in polynomial time. In this project, we will review quantum computing basics using real quantum computers and build on those basics to solve a subset of a graph optimization problems using both existing and …


Ur-269 Towards Bounding The Behavior Of Deep Neural Networks, Richard Borowski Dec 2022

Ur-269 Towards Bounding The Behavior Of Deep Neural Networks, Richard Borowski

C-Day Computing Showcase

Advances in Artificial Intelligence (AI), particularly in the form of deep neural networks, have revolutionized a diverse range of fields. As neural networks become more pervasive, the need to understand the boundaries of their behavior is becoming increasingly important. For example, can we formally guarantee that an autonomous vehicle will not violate traffic laws, such as reaching excessive speeds? Towards the goal of bounding the behavior of a neural network, we propose how to bound the behavior of individual neurons by incrementally tightening formal bounds on it. We further provide a case study on classifying handwritten digits to illustrate the …


Ur-294 A Quantum Arithmetic Logic Unit, Ethan Butler, Bryson Phillip, Benjamin G. Ulrich, David Carroll Dec 2022

Ur-294 A Quantum Arithmetic Logic Unit, Ethan Butler, Bryson Phillip, Benjamin G. Ulrich, David Carroll

C-Day Computing Showcase

This paper demonstrates that a quantum version of a classical Arithmetic Logic Unit (ALU) can be implemented on a quantum circuit. It would perform the same functions as the classical ALU, with the possibility of adding quantum functions in conjunction. To create the quantum ALU, we utilized IBM’s Qiskit Python package and JuypterLab. We also used the IBM Quantum Lab to run the circuit. We believe that a quantum ALU has the potential to be faster than its classical counterpart and the ability to calculate quantum specific operations. The simple classical functions translated to a quantum circuit show a promising …


Ur-295 Data Collection In Parkinson's Vr, Neil E. Weingarten, Ian Mcconnell Dec 2022

Ur-295 Data Collection In Parkinson's Vr, Neil E. Weingarten, Ian Mcconnell

C-Day Computing Showcase

This project is meant to show an addition to a Parkinson's simulation within VR where there are now different methods of data collection that are collected in-game. These data points are tracked and logged during gameplay, and are meant to allow researchers to make more effective use of the simulation as a tool for data collection. An example demo of the game and example files that were generated during gameplay are provided.


Predicting Publication Of Clinical Trials Using Structured And Unstructured Data: Model Development And Validation Study, Siyang Wang, Simon Šuster, Timothy Baldwin, Karin Verspoor Dec 2022

Predicting Publication Of Clinical Trials Using Structured And Unstructured Data: Model Development And Validation Study, Siyang Wang, Simon Šuster, Timothy Baldwin, Karin Verspoor

Natural Language Processing Faculty Publications

Background: Publication of registered clinical trials is a critical step in the timely dissemination of trial findings. However, a significant proportion of completed clinical trials are never published, motivating the need to analyze the factors behind success or failure to publish. This could inform study design, help regulatory decision-making, and improve resource allocation. It could also enhance our understanding of bias in the publication of trials and publication trends based on the research direction or strength of the findings. Although the publication of clinical trials has been addressed in several descriptive studies at an aggregate level, there is a lack …


Assisting The Human Fact-Checkers: Detecting All Previously Fact-Checked Claims In A Document, Shaden Shaar, Nikola Georgiev, Firoj Alam, Giovanni Da San Martino, Aisha Mohamed, Preslav Nakov Dec 2022

Assisting The Human Fact-Checkers: Detecting All Previously Fact-Checked Claims In A Document, Shaden Shaar, Nikola Georgiev, Firoj Alam, Giovanni Da San Martino, Aisha Mohamed, Preslav Nakov

Natural Language Processing Faculty Publications

Given the recent proliferation of false claims online, there has been a lot of manual fact-checking effort. As this is very time-consuming, human fact-checkers can benefit from tools that can support them and make them more efficient. Here, we focus on building a system that could provide such support. Given an input document, it aims to detect all sentences that contain a claim that can be verified by some previously fact-checked claims (from a given database). The output is a re-ranked list of the document sentences, so that those that can be verified are ranked as high as possible, together …


Overview Of The Wanlp 2022 Shared Task On Propaganda Detection In Arabic, Firoj Alam, Hamdy Mubarak, Wajdi Zaghouani, Giovanni Da San Martino, Preslav Nakov Dec 2022

Overview Of The Wanlp 2022 Shared Task On Propaganda Detection In Arabic, Firoj Alam, Hamdy Mubarak, Wajdi Zaghouani, Giovanni Da San Martino, Preslav Nakov

Natural Language Processing Faculty Publications

Propaganda is the expression of an opinion or an action by an individual or a group deliberately designed to influence the opinions or the actions of other individuals or groups with reference to predetermined ends, which is achieved by means of well-defined rhetorical and psychological devices. Propaganda techniques are commonly used in social media to manipulate or to mislead users. Thus, there has been a lot of recent research on automatic detection of propaganda techniques in text as well as in memes. However, so far the focus has been primarily on English. With the aim to bridge this language gap, …


Which Interval-Valued Alternatives Are Possibly Optimal If We Use Hurwicz Criterion, Marina Tuyako Mizukoshi, Weldon Lodwick, Martine Ceberio, Vladik Kreinovich Dec 2022

Which Interval-Valued Alternatives Are Possibly Optimal If We Use Hurwicz Criterion, Marina Tuyako Mizukoshi, Weldon Lodwick, Martine Ceberio, Vladik Kreinovich

Departmental Technical Reports (CS)

In many practical situations, for each alternative i, we do not know the corresponding gain xi, we only know the interval [li,ui] of possible gains. In such situations, a reasonable way to select an alternative is to choose some value α from the interval [0,1] and select the alternative i for which the Hurwicz combination α*ui + (1 − α)*li is the largest possible. In situations when we do not know the user's α, a reasonable idea is to select all alternatives that are optimal for some α. In this paper, we describe a feasible algorithm for such a selection.


Supervised Acoustic Embeddings And Their Transferability Across Languages, Sreepratha Ram, Hanan Aldarmaki Dec 2022

Supervised Acoustic Embeddings And Their Transferability Across Languages, Sreepratha Ram, Hanan Aldarmaki

Natural Language Processing Faculty Publications

In speech recognition, it is essential to model the phonetic content of the input signal while discarding irrelevant factors such as speaker variations and noise, which is challenging in low-resource settings. Self-supervised pretraining has been proposed as a way to improve both supervised and unsupervised speech recognition, including frame-level feature representations and Acoustic Word Embeddings (AWE) for variable-length segments. However, self-supervised models alone cannot learn perfect separation of the linguistic content as they are trained to optimize indirect objectives. In this work, we experiment with different pre-trained self-supervised features as input to AWE models and show that they work best …


(R1971) Analysis Of Feedback Queueing Model With Differentiated Vacations Under Classical Retrial Policy, Poonam Gupta, Naveen Kumar, Rajni Gupta Dec 2022

(R1971) Analysis Of Feedback Queueing Model With Differentiated Vacations Under Classical Retrial Policy, Poonam Gupta, Naveen Kumar, Rajni Gupta

Applications and Applied Mathematics: An International Journal (AAM)

This paper analyzes an M/M/1 retrial queue under differentiated vacations and Bernoulli feedback policy. On receiving the service, if the customer is not satisfied, then he may join the retrial group again with some probability and demand for service or may leave the system with the complementary probability. Using the probability generating functions technique, the steady-state solutions of the system are obtained. Furthermore, we have obtained some of the important performance measures such as expected orbit length, expected length of the system, sojourn times and probability of server being in different states. Using MATLAB software, we have represented the graphical …


(R1984) Analysis Of M^[X1], M^[X2]/G1, G_2^(A,B)/1 Queue With Priority Services, Server Breakdown, Repair, Modified Bernoulli Vacation, Immediate Feedback, G. Ayyappan, S. Nithya, B. Somasundaram Dec 2022

(R1984) Analysis Of M^[X1], M^[X2]/G1, G_2^(A,B)/1 Queue With Priority Services, Server Breakdown, Repair, Modified Bernoulli Vacation, Immediate Feedback, G. Ayyappan, S. Nithya, B. Somasundaram

Applications and Applied Mathematics: An International Journal (AAM)

In this investigation, the steady state analysis of two individualistic batch arrival queues with immediate feedback, modified Bernoulli vacation and server breakdown are introduced. Two different categories of customers like priority and ordinary are to be considered. This model propose nonpreemptive priority discipline. Ordinary and priority customers arrive as per Poisson processes. The server consistently afford single service for priority customers and the general bulk service for the ordinary customers and the service follows general distribution. The ordinary customers to be served only if the batch size should be greater than or equal to "a", else the server should not …


Camelira: An Arabic Multi-Dialect Morphological Disambiguator, Ossama Obeid, Go Inoue, Nizar Habash Dec 2022

Camelira: An Arabic Multi-Dialect Morphological Disambiguator, Ossama Obeid, Go Inoue, Nizar Habash

Computer Vision Faculty Publications

We present Camelira, a web-based Arabic multi-dialect morphological disambiguation tool that covers four major variants of Arabic: Modern Standard Arabic, Egyptian, Gulf, and Levantine. Camelira offers a user-friendly web interface that allows researchers and language learners to explore various linguistic information, such as part-of-speech, morphological features, and lemmas. Our system also provides an option to automatically choose an appropriate dialect-specific disambiguator based on the prediction of a dialect identification component. Camelira is publicly accessible at http://camelira.camel-lab.com.


Pasta: Table-Operations Aware Fact Verification Via Sentence-Table Cloze Pre-Training, Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, Xiaoyong Du Dec 2022

Pasta: Table-Operations Aware Fact Verification Via Sentence-Table Cloze Pre-Training, Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, Xiaoyong Du

Natural Language Processing Faculty Publications

Fact verification has attracted a lot of research attention recently, e.g., in journalism, marketing, and policymaking, as misinformation and disinformation online can sway one's opinion and affect one's actions. While fact-checking is a hard task in general, in many cases, false statements can be easily debunked based on analytics over tables with reliable information. Hence, table-based fact verification has recently emerged as an important and growing research area. Yet, progress has been limited due to the lack of datasets that can be used to pre-train language models (LMs) to be aware of common table operations, such as aggregating a column …


Greener: Graph Neural Networks For News Media Profiling, Panayot Panayotov, Utsav Shukla, Husrev T. Sencar, Mohamed Nabeel, Preslav Nakov Dec 2022

Greener: Graph Neural Networks For News Media Profiling, Panayot Panayotov, Utsav Shukla, Husrev T. Sencar, Mohamed Nabeel, Preslav Nakov

Natural Language Processing Faculty Publications

We study the problem of profiling news media on the Web with respect to their factuality of reporting and bias. This is an important but under-studied problem related to disinformation and “fake news” detection, but it addresses the issue at a coarser granularity compared to looking at an individual article or an individual claim. This is useful as it allows to profile entire media outlets in advance. Unlike previous work, which has focused primarily on text (e.g., on the articles published by the target website, or on the textual description in their social media profiles or in Wikipedia), here we …


Asdot: Any-Shot Data-To-Text Generation With Pretrained Language Models, Jiannan Xiang, Zhengzhong Liu, Yucheng Zhou, Eric P. Xing, Zhiting Hu Dec 2022

Asdot: Any-Shot Data-To-Text Generation With Pretrained Language Models, Jiannan Xiang, Zhengzhong Liu, Yucheng Zhou, Eric P. Xing, Zhiting Hu

Machine Learning Faculty Publications

Data-to-text generation is challenging due to the great variety of the input data in terms of domains (e.g., finance vs sports) or schemata (e.g., diverse predicates). Recent end-to-end neural methods thus require substantial training examples to learn to disambiguate and describe the data. Yet, real-world data-to-text problems often suffer from various data-scarce issues: one may have access to only a handful of or no training examples, and/or have to rely on examples in a different domain or schema. To fill this gap, we propose Any-Shot Data-to-Text (ASDOT), a new approach flexibly applicable to diverse settings by making efficient use of …


Impact Of Digital Twins And Metaverse On Cities: History, Current Situation, And Application Perspectives, Zhihan Lv, Wen Long Shang, Mohsen Guizani Dec 2022

Impact Of Digital Twins And Metaverse On Cities: History, Current Situation, And Application Perspectives, Zhihan Lv, Wen Long Shang, Mohsen Guizani

Machine Learning Faculty Publications

To promote the expansion and adoption of Digital Twins (DTs) in Smart Cities (SCs), a detailed review of the impact of DTs and digitalization on cities is made to assess the progression of cities and standardization of their management mode. Combined with the technical elements of DTs, the coupling effect of DTs technology and urban construction and the internal logic of DTs technology embedded in urban construction are discussed. Relevant literature covering the full range of DTs technologies and their applications is collected, evaluated, and collated, relevant studies are concatenated, and relevant accepted conclusions are summarized by modules. First, the …


Resel: N-Ary Relation Extraction From Scientific Text And Tables By Learning To Retrieve And Select, Yuchen Zhuang, Yinghao Li, Jerry Junyang Cheung, Yue Yu, Yingjun Mou, Xiang Chen, Le Song, Chao Zhang Dec 2022

Resel: N-Ary Relation Extraction From Scientific Text And Tables By Learning To Retrieve And Select, Yuchen Zhuang, Yinghao Li, Jerry Junyang Cheung, Yue Yu, Yingjun Mou, Xiang Chen, Le Song, Chao Zhang

Machine Learning Faculty Publications

We study the problem of extracting N-ary relation tuples from scientific articles. This task is challenging because the target knowledge tuples can reside in multiple parts and modalities of the document. Our proposed method RESEL decomposes this task into a two-stage procedure that first retrieves the most relevant paragraph/table and then selects the target entity from the retrieved component. For the high-level retrieval stage, RESEL designs a simple and effective feature set, which captures multilevel lexical and semantic similarities between the query and components. For the low-level selection stage, RESEL designs a cross-modal entity correlation graph along with a multi-view …


Efficient (Soft) Q-Learning For Text Generation With Limited Good Data, Han Guo, Bowen Tan, Zhengzhong Liu, Eric P. Xing, Zhiting Hu Dec 2022

Efficient (Soft) Q-Learning For Text Generation With Limited Good Data, Han Guo, Bowen Tan, Zhengzhong Liu, Eric P. Xing, Zhiting Hu

Machine Learning Faculty Publications

Maximum likelihood estimation (MLE) is the predominant algorithm for training text generation models. This paradigm relies on direct supervision examples, which is not applicable to many emerging applications, such as generating adversarial attacks or generating prompts to control language models. Reinforcement learning (RL) on the other hand offers a more flexible solution by allowing users to plug in arbitrary task metrics as reward. Yet previous RL algorithms for text generation, such as policy gradient (on-policy RL) and Q-learning (off-policy RL), are often notoriously inefficient or unstable to train due to the large sequence space and the sparse reward received only …


Amp: Automatically Finding Model Parallel Strategies With Heterogeneity Awareness, Dacheng Li, Hongyi Wang, Eric Xing, Hao Zhang Dec 2022

Amp: Automatically Finding Model Parallel Strategies With Heterogeneity Awareness, Dacheng Li, Hongyi Wang, Eric Xing, Hao Zhang

Machine Learning Faculty Publications

Scaling up model sizes can lead to fundamentally new capabilities in many machine learning (ML) tasks. However, training big models requires strong distributed system expertise to carefully design model-parallel execution strategies that suit the model architectures and cluster setups. In this paper, we develop AMP, a framework that automatically derives such strategies. AMP identifies a valid space of model parallelism strategies and efficiently searches the space for high-performed strategies, by leveraging a cost model designed to capture the heterogeneity of the model and cluster specifications. Unlike existing methods, AMP is specifically tailored to support complex models composed of uneven layers …


Unpaired Image-To-Image Translation With Density Changing Regularization, Shaoan Xie, Qirong Ho, Kun Zhang Dec 2022

Unpaired Image-To-Image Translation With Density Changing Regularization, Shaoan Xie, Qirong Ho, Kun Zhang

Machine Learning Faculty Publications

Unpaired image-to-image translation aims to translate an input image to another domain such that the output image looks like an image from another domain while important semantic information are preserved. Inferring the optimal mapping with unpaired data is impossible without making any assumptions. In this paper, we make a density changing assumption where image patches of high probability density should be mapped to patches of high probability density in another domain. Then we propose an efficient way to enforce this assumption: we train the flows as density estimators and penalize the variance of density changes. Despite its simplicity, our method …


On Pac Learning Halfspaces In Non-Interactive Local Privacy Model With Public Unlabeled Data, Jinyan Su, Jinhui Xu, Di Wang Dec 2022

On Pac Learning Halfspaces In Non-Interactive Local Privacy Model With Public Unlabeled Data, Jinyan Su, Jinhui Xu, Di Wang

Machine Learning Faculty Publications

In this paper, we study the problem of PAC learning halfspaces in the non-interactive local differential privacy model (NLDP). To breach the barrier of exponential sample complexity, previous results studied a relaxed setting where the server has access to some additional public but unlabeled data. We continue in this direction. Specifically, we consider the problem under the standard setting instead of the large margin setting studied before. Under different mild assumptions on the underlying data distribution, we propose two approaches that are based on the Massart noise model and self-supervised learning and show that it is possible to achieve sample …


A Cybersecurity Assessment Of Health Data Ecosystems, Michelle N. Halsey Dec 2022

A Cybersecurity Assessment Of Health Data Ecosystems, Michelle N. Halsey

Cyber Operations and Resilience Program Graduate Projects

This paper is an exploratory study that investigates data collected and used by health plans and reviews the laws and regulations governing this data to identify the gaps in protections and provide recommendations for eliminating these gaps. Health insurance companies collect a wide array of data about the people they insure, data that is often only peripherally relevant to the service these companies provide. The data environment currently consists of seven categories of data: personal health information, summary health information, personally identifiable information, financial information, professional information, biometric information, and lifestyle data or social indicators of health. Much of this …


A Damped Newton Method Achieves Global O(1/K2) And Local Quadratic Convergence Rate, Slavomír Hanzely, Dmitry Kamzolov, Dmitry Pasechnyuk, Alexander Gasnikov, Peter Richtárik, Martin Takáč Dec 2022

A Damped Newton Method Achieves Global O(1/K2) And Local Quadratic Convergence Rate, Slavomír Hanzely, Dmitry Kamzolov, Dmitry Pasechnyuk, Alexander Gasnikov, Peter Richtárik, Martin Takáč

Machine Learning Faculty Publications

In this paper, we present the first stepsize schedule for Newton method resulting in fast global and local convergence guarantees. In particular, a) we prove an O (1/k2) global rate, which matches the state-of-the-art global rate of cubically regularized Newton method of Polyak and Nesterov (2006) and of regularized Newton method of Mishchenko (2021) and Doikov and Nesterov (2021), b) we prove a local quadratic rate, which matches the best-known local rate of second-order methods, and c) our stepsize formula is simple, explicit, and does not require solving any subproblem. Our convergence proofs hold under affine-invariance assumptions closely related to …


Automs: Automatic Model Selection For Novelty Detection With Error Rate Control, Yifan Zhang, Haiyan Jiang, Haojie Ren, Changliang Zou, Dejing Dou Dec 2022

Automs: Automatic Model Selection For Novelty Detection With Error Rate Control, Yifan Zhang, Haiyan Jiang, Haojie Ren, Changliang Zou, Dejing Dou

Machine Learning Faculty Publications

Given an unsupervised novelty detection task on a new dataset, how can we automatically select a “best” detection model while simultaneously controlling the error rate of the best model? For novelty detection analysis, numerous detectors have been proposed to detect outliers on a new unseen dataset based on a score function trained on available clean data. However, due to the absence of labeled anomalous data for model evaluation and comparison, there is a lack of systematic approaches that are able to select the “best” model/detector (i.e., the algorithm as well as its hyperparameters) and achieve certain error rate control simultaneously. …