Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (9)
- Mathematics (6)
- Other Computer Sciences (6)
- Software Engineering (6)
- Numerical Analysis and Scientific Computing (4)
-
- Applied Mathematics (3)
- Data Science (3)
- Statistics and Probability (3)
- Arts and Humanities (2)
- Categorical Data Analysis (2)
- Discrete Mathematics and Combinatorics (2)
- Graphics and Human Computer Interfaces (2)
- Logic and Foundations (2)
- Systems Architecture (2)
- Algebra (1)
- Analysis (1)
- Applied Statistics (1)
- Biological and Chemical Physics (1)
- Computational Neuroscience (1)
- Computer Engineering (1)
- Computer and Systems Architecture (1)
- Controls and Control Theory (1)
- Data Storage Systems (1)
- Digital Communications and Networking (1)
- Electrical and Computer Engineering (1)
- Engineering (1)
- Information Security (1)
- Institution
- Publication Year
- Publication
- Publication Type
Articles 1 - 25 of 25
Full-Text Articles in Theory and Algorithms
Evaluation And Distillation Of Source Code Generation Tasks By Large Language Models, Danny Brahman
Evaluation And Distillation Of Source Code Generation Tasks By Large Language Models, Danny Brahman
Electronic Theses and Dissertations
Large Language Models (LLMs) are predominantly assessed based on their common sense reasoning, language comprehension, and logical reasoning abilities. While models trained in specialized domains like mathematics or coding have demonstrated remarkable advancements in logical reasoning, there remains a significant gap in evaluating their code generation capabilities. Existing benchmark datasets fall short in pinpointing specific strengths and weaknesses, impeding targeted enhancements in models’ reasoning abilities to synthesize code.
To bridge this gap, this thesis introduces two novel contributions: CodeEval and CodeQual. CodeEval is an innovative, pedagogical benchmarking method that mirrors the evaluation processes encountered in academic programming courses. It comprises …
Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun
Addressing The Problems Of Data Variations, Quality, And Scarcity In Training Deep Neural Networks, Jian Sun
Electronic Theses and Dissertations
The performance of deep neural networks (DNNs) is strongly influenced by the characteristics and quality of the underlying datasets. This Ph.D. dissertation addresses three pervasive data challenges-imbalance, quality degradation, and scarcity-that commonly hinder the effectiveness of DNNs in computer vision (CV) and natural language processing (NLP) applications.
Class imbalance remains one of the most frequent causes of degraded model generalization. While Focal Loss effectively mitigates inter-class imbalance by assigning higher weights to minority classes, it struggles with intra-class imbalance, particularly in video datasets where longer clips dominate feature representation. To address this, I implement and utilize …
Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay
Csc36000 - Modern Distributed Computing Assignment, Saptarashmi Bandyopadhyay
Open Educational Resources
This assignment covers standard performance metrics for Distributed Systems and the basics of Multiprocessing for CSC36000 - Modern Distributed Computing at the City College of New York CUNY. It is an interactive coding assignment intended to be executed in a Python notebook.
Multi-Label Classification Of Acoustic And Electronic Drum Sounds Using Machine Learning, Sean Perman
Multi-Label Classification Of Acoustic And Electronic Drum Sounds Using Machine Learning, Sean Perman
Electronic Theses and Dissertations
This paper presents a system for multi-class classification of drum sounds using audio signal processing and machine learning techniques. The project utilizes a diverse dataset of both acoustic and electronic drum samples and extracts ten distinct audio features to capture the timbral and temporal characteristics of each sound. The methodology includes signal preprocessing, feature extraction, and the application of supervised classification algorithms to distinguish between multiple drum classes. Experimental evaluations demonstrate that the selected features significantly enhance classification accuracy across a varied dataset. These findings underscore the effectiveness of combining traditional audio processing with modern machine learning, offering promising applications …
Optimizing Option Market Clearing, Juan Andrés Malaver Alvarado
Optimizing Option Market Clearing, Juan Andrés Malaver Alvarado
Electronic Theses and Dissertations
Modern options markets clear each strike in isolation, leaving cross-strike arbitrage unexploited. This thesis applies a payoff-dominant clearing mechanism to realized trades—roughly 2 000 Cboe VIX option executions from June–November 2016—after classifying each trade’s side and bundling by expiration. Three optimization formulations are tested: a fractional linear program (LP), a mixed-integer LP, and a pure integer program. On a 10-core laptop every bundle solves in < 0.5 s. The LP captures the greatest surplus, yet the integer models recover nearly as much while filling whole contracts and holding only modest margin. Results reveal persistent, albeit small, inefficiencies in executed trades and demonstrate that an integral cross-strike auction could operate in real time. The accompanying C/Gurobi code is modular and readily extendable to early-exercise options. Trade-level evidence thus supports redesigning exchange clearing to consider the complete option book.
Investigating Key Structures In Protective Scenes For Llms, Eben M. Weisman
Investigating Key Structures In Protective Scenes For Llms, Eben M. Weisman
University Honors Theses
This research delves into the realm of "protective scenes" within Large Language Models (LLMs), exploring their impact on bias mitigation, deception, and context preservation. The study investigates the use of roleplay prompting human-like behavior and reasoning in LLMs, focusing on the Character-LLM framework's concept of protective scenes with graduated levels of protection. By combining insights from psychology, cognitive science, and computational analysis, this research aims to develop a framework for understanding how protective scenes influence roleplay performance in LLMs, ultimately contributing to the development of more reliable and ethical AI systems.
Towards Erasing The Distinction Between The Computational And Syntactic Accounts Of Scientific Theories, Timothy Luft
Towards Erasing The Distinction Between The Computational And Syntactic Accounts Of Scientific Theories, Timothy Luft
Theses
One of the main goals of philosophy of science is to give a proper account of scientific theories and their structure. One way that accounts of the structure of scientific theories can be distinguished is by the mathematical or logical structures that they involve. For instance, syntactic accounts of scientific theories hold that theories are axioms in a logical framework, whereas semantic accounts are more liberal in the range of mathematical and logical structures they take as pertinent to the structure of scientific theories. Paul Thagard (1988) offers a computational account of scientific theories, which holds that theories are complex …
Predicting Biomolecular Properties And Interactions Using Numerical, Statistical And Machine Learning Methods, Elyssa Sliheet
Predicting Biomolecular Properties And Interactions Using Numerical, Statistical And Machine Learning Methods, Elyssa Sliheet
Mathematics Theses and Dissertations
We investigate machine learning and electrostatic methods to predict biophysical properties of proteins, such as solvation energy and protein ligand binding affinity, for the purpose of drug discovery/development. We focus on the Poisson-Boltzmann model and various high performance computing considerations such as parallelization schemes.
An Empirical Study Of Machine Learning Techniques For Accurate Stock Price Forecasting, Daniel Paliulis, Hari Patchigolla
An Empirical Study Of Machine Learning Techniques For Accurate Stock Price Forecasting, Daniel Paliulis, Hari Patchigolla
Honors Scholar Theses
This paper presents a comprehensive approach to predicting future stock prices of companies using machine learning and time series analysis. The research problem is centered around addressing the complexity and emotion-driven nature of stock investment decisions. To create an objective determinant in stock decisions, we propose a machine learning model utilizing time series data from major companies, including Amazon, Apple, Google, Nvidia, Meta, Tesla, Salesforce, Intel, and Microsoft. We explore the use of Long Short-Term Memory (LSTM) neural networks, to capture the temporal dynamics of stock prices. These models are designed to process sequential data, maintaining short term and long …
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi
Enhancing Search Engine Results: A Comparative Study Of Graph And Timeline Visualizations For Semantic And Temporal Relationship Discovery, Muhammad Shahiq Qureshi
Electronic Theses and Dissertations
In today’s digital age, search engines have become indispensable tools for finding information among the corpus of billions of webpages. The standard that most search engines follow is to display search results in a list-based format arranged according to a ranking algorithm. Although this format is good for presenting the most relevant results to users, it fails to represent the underlying relations between different results. These relations, among others, can generally be of either a temporal or semantic nature. A user who wants to explore the results that are connected by those relations would have to make a manual effort …
Visualized Algorithm Engineering On Two Graph Partitioning Problems, Zizhen Chen
Visualized Algorithm Engineering On Two Graph Partitioning Problems, Zizhen Chen
Computer Science and Engineering Theses and Dissertations
Concepts of graph theory are frequently used by computer scientists as abstractions when modeling a problem. Partitioning a graph (or a network) into smaller parts is one of the fundamental algorithmic operations that plays a key role in classifying and clustering. Since the early 1970s, graph partitioning rapidly expanded for applications in wide areas. It applies in both engineering applications, as well as research. Current technology generates massive data (“Big Data”) from business interactions and social exchanges, so high-performance algorithms of partitioning graphs are a critical need.
This dissertation presents engineering models for two graph partitioning problems arising from completely …
Procedural Level Generation For A Top-Down Roguelike Game, Kieran Ahn, Tyler Edmiston
Procedural Level Generation For A Top-Down Roguelike Game, Kieran Ahn, Tyler Edmiston
Honors Thesis
In this file, I present a sequence of algorithms that handle procedural level generation for the game Fragment, a game designed for CMSI 4071 and CMSI 4071 in collaboration with students from the LMU Animation department. I use algorithms inspired by graph theory and implementing best practices to the best of my ability. The full level generation sequence is comprised of four algorithms: the terrain generation, boss room placement, player spawn point selection, and enemy population. The terrain generation algorithm takes advantage of tree traversal methods to create a connected graph of walkable tiles. The boss room placement algorithm randomly …
Crosshair Optimizer, Jason Torrence
Crosshair Optimizer, Jason Torrence
All Master's Theses
Metaheuristic optimization algorithms are heuristics that are capable of creating a "good enough'' solution to a computationally complex problem. Algorithms in this area of study are focused on the process of exploration and exploitation: exploration of the solution space and exploitation of the results that have been found during that exploration, with most resources going toward the former half of the process. The novel Crosshair optimizer developed in this thesis seeks to take advantage of the latter, exploiting the best possible result as much as possible by directly searching the area around that best result with a stochastic approach. This …
The “Knapsack Problem” Workbook: An Exploration Of Topics In Computer Science, Steven Cosares
The “Knapsack Problem” Workbook: An Exploration Of Topics In Computer Science, Steven Cosares
Open Educational Resources
This workbook provides discussions, programming assignments, projects, and class exercises revolving around the “Knapsack Problem” (KP), which is widely a recognized model that is taught within a typical Computer Science curriculum. Throughout these discussions, we use KP to introduce or review topics found in courses covering topics in Discrete Mathematics, Mathematical Programming, Data Structures, Algorithms, Computational Complexity, etc. Because of the broad range of subjects discussed, this workbook and the accompanying spreadsheet files might be used as part of some CS capstone experience. Otherwise, we recommend that individual sections be used, as needed, for exercises relevant to a course in …
Analysis Of Github Pull Requests, Canon Ellis
Analysis Of Github Pull Requests, Canon Ellis
Computer Science and Engineering Theses and Dissertations
The popularity of the software repository site GitHub has created a rise in the Pull Based Development Models' use. An essential portion of pull-based development is the creation of Pull Requests. Pull Requests often have to be reviewed by an individual to be approved and accepted into the Master branch of a software repository. The reviewing process can often be time-consuming and introduce a relatively high level of lost development time. This paper examines thousands of pull requests to understand the most valuable metadata of pull requests. We then introduce metrics in comparing the metadata of pull requests to understand …
Heuristics For Sparsest Cut Approximations In Network Flow Applications, Fernando Vilas
Heuristics For Sparsest Cut Approximations In Network Flow Applications, Fernando Vilas
Computer Science and Engineering Theses and Dissertations
The Maximum Concurrent Flow Problem (MCFP) is a polynomially bounded problem that has been used over the years in a variety of applications. Sometimes it is used to attempt to find the Sparsest Cut, an NP-hard problem, and other times to find communities in Social Network Analysis (SNA) in its hierarchical formulation, the HMCFP. Though it is polynomially bounded, the MCFP quickly grows in space utilization, rendering it useful on only small problems. When it was defined, only a few hundred nodes could be solved, where a few decades later, graphs of one to two thousand nodes can still be …
Using Natural Language Processing To Categorize Fictional Literature In An Unsupervised Manner, Dalton J. Crutchfield
Using Natural Language Processing To Categorize Fictional Literature In An Unsupervised Manner, Dalton J. Crutchfield
Electronic Theses and Dissertations
When following a plot in a story, categorization is something that humans do without even thinking; whether this is simple classification like “This is science fiction” or more complex trope recognition like recognizing a Chekhov's gun or a rags to riches storyline, humans group stories with other similar stories. Research has been done to categorize basic plots and acknowledge common story tropes on the literary side, however, there is not a formula or set way to determine these plots in a story line automatically. This paper explores multiple natural language processing techniques in an attempt to automatically compare and cluster …
Stochastic Orthogonalization And Its Application To Machine Learning, Yu Hong
Stochastic Orthogonalization And Its Application To Machine Learning, Yu Hong
Electrical Engineering Theses and Dissertations
Orthogonal transformations have driven many great achievements in signal processing. They simplify computation and stabilize convergence during parameter training. Researchers have introduced orthogonality to machine learning recently and have obtained some encouraging results. In this thesis, three new orthogonal constraint algorithms based on a stochastic version of an SVD-based cost are proposed, which are suited to training large-scale matrices in convolutional neural networks. We have observed better performance in comparison with other orthogonal algorithms for convolutional neural networks.
Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane
Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane
Statistical Science Theses and Dissertations
If the Warriors beat the Rockets and the Rockets beat the Spurs, does that mean that the Warriors are better than the Spurs? Sophisticated fans would argue that the Warriors are better by the transitive property, but could Spurs fans make a legitimate argument that their team is better despite this chain of evidence?
We first explore the nature of intransitive (rock-scissors-paper) relationships with a graph theoretic approach to the method of paired comparisons framework popularized by Kendall and Smith (1940). Then, we focus on the setting where all pairs of items, teams, players, or objects have been compared to …
A Hott Approach To Computational Effects, Phillip A. Wells
A Hott Approach To Computational Effects, Phillip A. Wells
Senior Independent Study Theses
A computational effect is any mutation of real-world state that occurs as the result of a computation. We develop a model for describing computational effects within homotopy type theory, a branch of mathematics separate from other foundations such as set theory. Such a model allows us to describe programs as total functions over values while preserving information about the effects those programs induce.
Understanding Natural Keyboard Typing Using Convolutional Neural Networks On Mobile Sensor Data, Travis Siems
Understanding Natural Keyboard Typing Using Convolutional Neural Networks On Mobile Sensor Data, Travis Siems
Computer Science and Engineering Theses and Dissertations
Mobile phones and other devices with embedded sensors are becoming increasingly ubiquitous. Audio and motion sensor data may be able to detect information that we did not think possible. Some researchers have created models that can predict computer keyboard typing from a nearby mobile device; however, certain limitations to their experiment setup and methods compelled us to be skeptical of the models’ realistic prediction capability. We investigate the possibility of understanding natural keyboard typing from mobile phones by performing a well-designed data collection experiment that encourages natural typing and interactions. This data collection helps capture realistic vulnerabilities of the security …
Logic -> Proof -> Rest, Maxwell Taylor
Logic -> Proof -> Rest, Maxwell Taylor
Senior Independent Study Theses
REST is a common architecture for networked applications. Applications that adhere to the REST constraints enjoy significant scaling advantages over other architectures. But REST is not a panacea for the task of building correct software. Algebraic models of computation, particularly CSP, prove useful to describe the composition of applications using REST. CSP enables us to describe and verify the behavior of RESTful systems. The descriptions of each component can be used independently to verify that a system behaves as expected. This thesis demonstrates and develops CSP methodology to verify the behavior of RESTful applications.
Exploring Algorithmic Musical Key Recognition, Nathan J. Levine
Exploring Algorithmic Musical Key Recognition, Nathan J. Levine
CMC Senior Theses
The following thesis outlines the goal and process of algorithmic musical key detection as well as the underlying music theory. This includes a discussion of signal-processing techniques intended to most accurately detect musical pitch, as well as a detailed description of the Krumhansl-Shmuckler (KS) key-finding algorithm. It also describes the Java based implementation and testing process of a musical key-finding program based on the KS algorithm. This thesis provides an analysis of the results and a comparison with the original algorithm, ending with a discussion of the recommended direction of further development.
Learning Emotions: A Software Engine For Simulating Realistic Emotion In Artificial Agents, Douglas Code
Learning Emotions: A Software Engine For Simulating Realistic Emotion In Artificial Agents, Douglas Code
Senior Independent Study Theses
This paper outlines a software framework for the simulation of dynamic emotions in simulated agents. This framework acts as a domain-independent, black-box solution for giving actors in games or simulations realistic emotional reactions to events. The emotion management engine provided by the framework uses a modified Fuzzy Logic Adaptive Model of Emotions (FLAME) model, which lets it manage both appraisal of events in relation to an individual’s emotional state, and learning mechanisms through which an individual’s emotional responses to a particular event or object can change over time. In addition to the FLAME model, the engine draws on the design …
Random Number Generation: Types And Techniques, David F. Dicarlo
Random Number Generation: Types And Techniques, David F. Dicarlo
Senior Honors Theses
What does it mean to have random numbers? Without understanding where a group of numbers came from, it is impossible to know if they were randomly generated. However, common sense claims that if the process to generate these numbers is truly understood, then the numbers could not be random. Methods that are able to let their internal workings be known without sacrificing random results are what this paper sets out to describe. Beginning with a study of what it really means for something to be random, this paper dives into the topic of random number generators and summarizes the key …