Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (23)
- Statistics and Probability (14)
- Artificial Intelligence and Robotics (9)
- Databases and Information Systems (9)
- Mathematics (7)
-
- Business (6)
- Social and Behavioral Sciences (6)
- Statistical Models (5)
- Applied Mathematics (4)
- Engineering (4)
- Life Sciences (4)
- Medicine and Health Sciences (4)
- Theory and Algorithms (4)
- Computer Engineering (3)
- Education (3)
- Information Security (3)
- Neuroscience and Neurobiology (3)
- Other Computer Sciences (3)
- Analysis (2)
- Applied Statistics (2)
- Arts and Humanities (2)
- Categorical Data Analysis (2)
- Chemistry (2)
- Computational Neuroscience (2)
- Educational Leadership (2)
- Finance and Financial Management (2)
- Numerical Analysis and Computation (2)
- Numerical Analysis and Scientific Computing (2)
- Institution
-
- Southern Methodist University (14)
- California Polytechnic State University, San Luis Obispo (4)
- Kennesaw State University (3)
- Belmont University (2)
- City University of New York (CUNY) (2)
-
- Claremont Colleges (2)
- Dartmouth College (2)
- The University of Akron (2)
- University of Arkansas, Fayetteville (2)
- University of New Mexico (2)
- University of South Carolina (2)
- Arcadia University (1)
- Bellarmine University (1)
- Bowling Green State University (1)
- California State University, San Bernardino (1)
- Clemson University (1)
- LSU New Orleans (1)
- Louisiana State University (1)
- Mississippi State University (1)
- Murray State University (1)
- National Louis University (1)
- Old Dominion University (1)
- Seattle Pacific University (1)
- Smith College (1)
- Technological University Dublin (1)
- University of Arkansas Little Rock (1)
- University of Central Florida (1)
- University of Connecticut (1)
- University of Mary Washington (1)
- University of New Hampshire (1)
- Publication
-
- SMU Data Science Review (14)
- Master's Theses (4)
- CMC Senior Theses (2)
- Dartmouth College Undergraduate Theses (2)
- Dissertations (2)
-
- Honors Projects (2)
- SPARK Symposium Presentations (2)
- Senior Theses (2)
- Williams Honors College, Honors Research Projects (2)
- All Dissertations (1)
- Articles (1)
- BCoE Publications (1)
- Capstone Showcase (1)
- Chemistry and Chemical Biology ETDs (1)
- Computer Science ETDs (1)
- Computer Science Theses & Dissertations (1)
- Computer Science and Computer Engineering Undergraduate Honors Theses (1)
- Data Science Undergraduate Honors Theses (1)
- Departmental Honors & Graduate Capstone Projects (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctor of Data Science and Analytics Dissertations (1)
- Electronic Theses, Projects, and Dissertations (1)
- Faculty Publications (1)
- Honors College Theses (1)
- Honors Scholar Theses (1)
- Honors Undergraduate Theses (1)
- LSU Master's Theses (1)
- LSU New Orleans Theses and Dissertations (1)
- Publications and Research (1)
- Published and Grey Literature from PhD Candidates (1)
- Publication Type
- File Type
Articles 1 - 30 of 57
Full-Text Articles in Data Science
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Master's Theses
Unsupervised clustering often faces the challenge of determining the correct number of clusters in the absence of a true target variable. Traditional methods such as the Elbow Method and the Silhouette Score can produce ambiguous results and rely on assumptions about cluster shape or separation. To address this, we created the Clustering Rivals and Buddies (CRAB) algorithm which evaluates clusters based on stability across multiple subsamples. CRAB uses pairwise classifications to identify points that consistently group together called “Buddies” and points that remain separated called “Rivals.” Applied with K-means, CRAB accurately recovers underlying cluster structures in both spherical and non-spherical …
A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout
A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout
Data Science Undergraduate Honors Theses
The purpose of this paper is to analyze patterns between public safety and streetlighting for the City of Sugar Land, TX so that they may better protect their citizens. The data involved come from the City of Sugar Land’s public works division and include type and location for all the attributes. The method of doing so involved visualizing the patterns of streetlights and their closest light readings to visualize which streetlights are underperforming using the Shiny package in R. Statistical tests were also used to quantify the association between lighting, crime occurrence, and crosswalks. From this, and the literature review, …
Curriculum For A Two Semester Calculus Course Specializing In Life Science And Data Science, Patrick Mcclain
Curriculum For A Two Semester Calculus Course Specializing In Life Science And Data Science, Patrick Mcclain
LSU Master's Theses
Traditionally, introductory calculus has been designed for engineering and physics students, often leaving students majoring in data science or life sciences with a curriculum that lacks professional relevance and is overly reliant on problems that focus on computational fluency. This thesis proposes a two-semester sequence, called MATH 153X and MATH 154X, specifically tailored for the Louisiana State University (LSU) Dual Enrollment program and university-level data science and life sciences majors. By integrating modern computational tools—such as symbolic calculators and artificial intelligence (AI) tools—the proposed curriculum shifts the pedagogical focus from procedural symbolic manipulation toward conceptual literacy.
Through a series …
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
Ai-Powered Reporting For Improved Hospital Efficiency, Joel C. Laskow, Chris Papesh, Srishti Awasthi, Jacquelyn Cheun
SMU Data Science Review
This study explores the feasibility of an AI-powered chatbot for HIPAA-aligned intake of emergency room patients seeking treatment for overdose and violence. The system utilizes AWS Amplify, an encrypted EC2 instance, and a secure S3 Bucket house on Amazon Web Services. Chat functionality is powered by a multi-agentic framework operating on Anthropic’s Claude Sonnet 4. Manual evaluation and exact match testing reveal the system reliably obtains and records relevant information during intake. Future work will focus on expanding accessibility by integrating voice functionality, obtaining HIPAA compliance certifications, and incorporating the chat system into existing healthcare networks.
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
A Study Of Machine Learning Techniques In Solving Biochemical And Chemical Problems, Kenneth Micheal Plackowski
A Study Of Machine Learning Techniques In Solving Biochemical And Chemical Problems, Kenneth Micheal Plackowski
Chemistry and Chemical Biology ETDs
Data-driven approaches to solving problems in biology and chemistry require utilization of reliable techniques and machine learning algorithms are the modern reliable approach. This work presents three problems that involve use of supervised learning techniques when classification is the goal and unsupervised learning techniques when global data representation is the goal.
In the first problem, we demonstrate the use of unsupervised clustering techniques, self-organizing maps and K-means, to ascertain analyte detection capabilities of carbon nitride dots. In the second problem, we add scalability features to a functional group classification model applied to infrared data and evaluate its ability to inform …
Frequent Itemset Mining With Tidyclust In R, Andrew D. Kerr
Frequent Itemset Mining With Tidyclust In R, Andrew D. Kerr
Master's Theses
Unsupervised learning is closely associated with clustering, however other methods fall under this umbrella such as data mining. In R, the tidyclust package provides a unified interface for clustering models, yet lacks support for data mining. This thesis addresses this gap by introducing the Apriori and ECLAT algorithms into tidyclust, with a focus on frequent itemset mining. Unlike traditional clustering models, frequent itemsets produce groupings of column variables, rather than cluster labels or partitions of observations. To address this, a novel clustering approach is proposed: items (columns) are grouped based on their ”dominant” frequent itemset. A key contribution is a …
A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer
A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer
Computer Science ETDs
Modern drug discovery and chemical biology research relies heavily on analyzing bioassay data. One of the many challenges in bioassay data analysis is identifying false trails, i.e., chemical compounds which initially appear to have desirable activity but are found to be problematic upon further investigation. Badapple (the BioAssay-Data Associative Promiscuity Pattern Learning Engine) was created over ten years ago to help researchers identify promiscuous compounds and thus avoid a common source of these false trails. Through an effort involving software engineering, cheminformatics, and biomedical data science we have developed Badapple 2.0, which incorporates updated assay records and expanded data semantics. …
Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan
Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan
SMU Data Science Review
This paper presents an innovative approach to enhancing network security by integrating machine learning algorithms with fine-tuned large language models (LLMs) to provide an expert assistant querying. The proposed method utilizes machine learning for efficient preprocessing and feature extraction from log data, followed by the application of a fine-tuned LLM to analyze and interpret anomalies with greater accuracy. This dual-layer detection system is designed to improve the identification of subtle and sophisticated security threats. The research team’s extensive evaluation using real-world log datasets indicates that the combined approach increases detection rates and communicates results in an understandable manner, demonstrating its …
From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie
From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie
Undergraduate Theses
Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …
Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt
Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt
SPARK Symposium Presentations
AI text generation is rapidly developing, and, as a result, it is becoming increasingly difficult to differentiate it from human written text. Our base study by Leon Fröhling et al. proposed a feature-based detection model trained on GPT2, GPT3, and Grover data, as well as human-generated text. Our work extends their research by training a modified model with four neural networks on word embeddings, select features from the original study, as well as updated data (GPT3, GPT4, and Grover).
"Data Science For Digital Privacy: A Practical Guide For Non-Technical Audiences", Kayla Ahrndt
"Data Science For Digital Privacy: A Practical Guide For Non-Technical Audiences", Kayla Ahrndt
SPARK Symposium Presentations
As companies increasingly rely on consumer data for personalization and profit, the need for stronger user protections, security measures, and transparency in data practices grows. Legal frameworks must be continuously re-evaluated and updated to ensure accountability, while individuals must be equipped with the knowledge to make informed decisions about their digital presence. However, personal data privacy education remains widely inaccessible due to the technical language and the effort required to navigate complex policies. This project, presented in both zine and blog formats, addresses this gap by providing clear, actionable, and accessible recommendations for data privacy and personal cybersecurity. As an …
Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin
Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin
Williams Honors College, Honors Research Projects
Heart disease remains the leading cause of death in the United States, particularly among the elderly population. The growing availability of large-scale health data and the advancement of machine learning tools present an opportunity to create more accurate and individualized predictive models. This study utilizes a subset of the 2020 Behavioral Risk Factor Surveillance System (BRFSS) dataset, focusing on individuals aged 70 and above, to explore predictive modeling using logistic regression, random forests, and XGBoost. The models were evaluated using key performance metrics, including sensitivity, specificity, accuracy, and the area under the ROC curve (AUC). The findings suggest that while …
Theoretical Analysis Of Cnns For Automatic Seizure Detection In Eeg Signals, Jackson T. Small
Theoretical Analysis Of Cnns For Automatic Seizure Detection In Eeg Signals, Jackson T. Small
Honors Undergraduate Theses
Epilepsy is a common brain disorder where neurons in the brain rapidly fire, causing recurring seizures. The brain activity during a seizure can be detected by electroencephalogram (EEG) signals; however, this process is not only labor-intensive and time-consuming but is also subject to inter-rater variability, with a study showing only moderate agreement when diagnosing patients, even among experts. Convolutional Neural Networks (CNNs) are often proposed to detect seizures automatically, achieving high performance. The focus on performance comes at a cost of losing interpretability, leaving the model as effective but seen as a ’black box’. This thesis confronts the interpretability knowledge …
Competitive Conquest: Charting The Climb To Pokémon Supremacy, Robert Dilworth
Competitive Conquest: Charting The Climb To Pokémon Supremacy, Robert Dilworth
BCoE Publications
This manuscript presents a comprehensive exploration of optimizing Pokémon gameplay through data-driven methodologies, aimed at enhancing competitive performance in high-stakes environments. In the first section, we introduce a robust Pokémon teambuilding algorithm that leverages statistical analysis of championship-winning compositions. By employing multiple linear regression techniques, we predict team performance based on critical factors such as Base Stat Totals (BSTs) and various coverage types. This integration of data science principles into Pokémon strategy underscores the importance of offensive capabilities over defensive considerations, ultimately contributing to advancements in teambuilding strategies. Our proficiency in R programming facilitated the development of an efficient codebase …
Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller
Enhancing Imputation Accuracy: A Multi-Faceted Approach For Missing Data In Chicago Arrest Records, Steve Bramhall, Jae Chung, Nicholas Mueller
SMU Data Science Review
This paper introduces a novel approach to enhance the imputation process for missing data, utilizing crime records from Chicago with arrests as the target feature. Robust imputation techniques are crucial in the era of burgeoning datasets for generating reliable insights. Our core objective is to present an innovative method that improves imputation techniques, augmenting model performance and bolstering the reliability of analytical outcomes. Leveraging numeric crime data, we establish a Gradient Boosting (GBM) baseline model, then introduce ensemble methods including Random Forest and Decision Trees for further refinement. By systematically exploring multiple imputation processes, we establish a baseline for comparative …
Enhancing Shap With Multi-Core Parallelization And Distributed Computation, Matthew David, William Jones, Hayley Horn
Enhancing Shap With Multi-Core Parallelization And Distributed Computation, Matthew David, William Jones, Hayley Horn
SMU Data Science Review
In recent years, the adoption of complex machine learning algorithms, often perceived as “black box” models, has grown exponentially across various disciplines. However, the lack of understanding regarding how these models come to their predictions often fosters skepticism and mistrust. In response to the demand for transparency and interpretability, Explainable AI techniques, such as SHapley Additive exPlanations (SHAP), have emerged as powerful tools for comprehending and trusting these algorithms. However, SHAP has an exponential computational demand O( x2 ), where x is the number of features. This becomes increasingly problematic with the larger datasets standard in most industries. Many frameworks …
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
All Dissertations
The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …
Streaminghub - A Realtime Biosignal Processing Framework For Lab Scale Experimentation, Yasith Jayawardana
Streaminghub - A Realtime Biosignal Processing Framework For Lab Scale Experimentation, Yasith Jayawardana
Computer Science Theses & Dissertations
In human subjects research, biosignals such as eye movements, heart rate, and brain activity, are often collected and analyzed to find patterns with tangible real-world implications. Modern advancements in technology have sparked interest towards analyzing biosignals in realtime. When developing such algorithms, one may expect to find free, open-source tools that provide easy access to live, recorded, and simulated data streams. Yet, biosignal interfaces are often vendor-specific, making cross-vendor biosignal streaming non-trivial. Likewise, reading biosignal datasets is also non-trivial, as their content may be arranged quite differently.
To combat this divide, we provide the scientific community with a realtime biosignal …
Building Effective Large Language Model Agents, Sydney Holder, Shreyash Taywade
Building Effective Large Language Model Agents, Sydney Holder, Shreyash Taywade
SMU Data Science Review
The advancement of large language models (LLMs) has significantly expanded the influence of artificial intelligence across various sectors. This paper explores building LLM agents to power applications and examines what is necessary to build an efficient and helpful AI assistant. The research investigates the core components necessary to create specialized agents, facilitate collaboration in problem-solving, and improve human task performance. The development and application of tools designed to augment the capabilities of LLM agents are also explored. The paper addresses the potential risks of the unknowns, such as hallucinations, which can compromise the success of agent-based solutions within LLM applications. …
The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi
The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi
Computer Science and Computer Engineering Undergraduate Honors Theses
The strategic planning of offensive passing plays in the NFL incorporates numerous variables, including defensive coverages, player positioning, historical data, etc. This project develops an application using an analytical framework and an interactive model to simulate and visualize an NFL offense's passing strategy under varying conditions. Using R-programming and data management, the model dynamically represents potential passing routes in response to different defensive schemes. The system architecture integrates data from historical NFL league years to generate quantified route scores through designed mathematical equations. This allows for the prediction of potential passing routes for offensive skill players in response to the …
Sports Science: An Entrepreneurial Venture, Nicole J. Jones
Sports Science: An Entrepreneurial Venture, Nicole J. Jones
Senior Honors Projects
In sports science, ensuring maximum athlete safety and optimizing data utilization are pivotal yet leave room for further work. My project, Unbeaten SafeWare, addresses these critical issues by focusing on two primary concerns: preventing heat-related and cardiac illnesses, which are significant causes of athlete fatalities, and enhancing the transparency and utility of sports data. This initiative involves developing a shirt integrated with sensors to monitor vital signs and an athlete management system to handle data input, storage, analysis, and accessibility for athletes.
The project has advanced through the efforts of a multidisciplinary team, which includes biomedical engineering undergraduates, two faculty …
A Holistic And Collaborative Behavioral Health Detection Framework Using Sensitive Police Narratives, Martin Keagan Wynne Brown
A Holistic And Collaborative Behavioral Health Detection Framework Using Sensitive Police Narratives, Martin Keagan Wynne Brown
Dissertations
Identifying behavioral health is paramount for law enforcement officers to provide appropriate follow-up community care. In the current practice, law enforcement offices manually identify these behavioral health cases to allow the designation of the relevant follow-up resources. Police reports generated by officers' response to 911 calls remain an untapped resource for identifying such incidents. Therefore, we advocate for the incorporation of manual annotations from experts, natural language processing (NLP), active learning, advanced machine learning, and ensemble techniques to detect behavioral health cases within police reports. In this dissertation, we develop tools and frameworks to automatically detect behavioral health cases from …
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research, Musa M. Azeem
Elevating Academic Administration: A Comprehensive Faculty Dashboard For Tracking Student Evaluations And Research, Musa M. Azeem
Senior Theses
The USC Faculty Dashboard is a web application designed to revolutionize how department heads, professors, and instructors monitor progress and make decisions, providing a centralized hub for efficient data storage and analysis. Currently, there’s a gap in tools tailored for department heads to concisely manage the performance of their department, which our platform aims to fill. The USC Faculty Dashboard offers easy access to upload and view student evaluation and research information, empowering department heads to evaluate the performance of faculty members and seamlessly track their research grants, publications, and expenditures. Furthermore, professors and instructors gain personalized performance analysis tools, …
An Empirical Study Of Machine Learning Techniques For Accurate Stock Price Forecasting, Daniel Paliulis, Hari Patchigolla
An Empirical Study Of Machine Learning Techniques For Accurate Stock Price Forecasting, Daniel Paliulis, Hari Patchigolla
Honors Scholar Theses
This paper presents a comprehensive approach to predicting future stock prices of companies using machine learning and time series analysis. The research problem is centered around addressing the complexity and emotion-driven nature of stock investment decisions. To create an objective determinant in stock decisions, we propose a machine learning model utilizing time series data from major companies, including Amazon, Apple, Google, Nvidia, Meta, Tesla, Salesforce, Intel, and Microsoft. We explore the use of Long Short-Term Memory (LSTM) neural networks, to capture the temporal dynamics of stock prices. These models are designed to process sequential data, maintaining short term and long …
Investigation Into A Practical Application Of Reinforcement Learning For The Stock Market, Philip Traxler, Sadik Aman, Will Rogers, Allyn Okun
Investigation Into A Practical Application Of Reinforcement Learning For The Stock Market, Philip Traxler, Sadik Aman, Will Rogers, Allyn Okun
SMU Data Science Review
A major problem of the financial industry is the ability to adapt their trading strategies at the same rate the market evolves. This paper proposes a solution using existing Reinforcement Learning libraries to help find new strategies at a practical scale. Using a wide domain of ticker symbols, an algorithm is trained in an environment that better represents reality. The supplied decision-making algorithm is tested using recorded data from the U.S stock market from 2000 through 2022. The results of this research show that existing techniques are statistically better than making decisions at random. With this result, this research shows …
A Prompt Engineering Approach To Creating Automated Commentary For Microsoft Self-Help Documentation Metric Reports Using Chatgpt, Ryan Herrin, Luke Stodgel, Brian Raffety
A Prompt Engineering Approach To Creating Automated Commentary For Microsoft Self-Help Documentation Metric Reports Using Chatgpt, Ryan Herrin, Luke Stodgel, Brian Raffety
SMU Data Science Review
Microsoft collects an immense amount of data from the users of their product-self-help documentation. Employees use this data to identify these self-help articles' performance trends and measure their impact on business Key Performance Indicators (KPIs). Microsoft uses various tools like Power BI and Python to analyze this data. The problem is that their analysis and findings are summarized manually. Therefore, this research will improve upon their current analysis methods by applying the latest prompt engineering practices and the power of ChatGPT's large language models (LLMs). Using VBA code, Microsoft Excel, and the ChatGPT API as an Excel add-in, this research …
General Population Projection Model With Census Population Data, Takenori Tsuruga
General Population Projection Model With Census Population Data, Takenori Tsuruga
Electronic Theses, Projects, and Dissertations
The US Census Bureau offers a wide range of data, and within this array, the American Community Survey 5-Year Estimate (ACS5) serves as a valuable resource for understanding the US population. This project embarks on an exploration of Machine Learning and the Software Development process with the goal of generating effective population projections from ACS5 data. The project aims to provide methods to make predictions for every city and town in the US, encompassing their total population and population divided into 5-year age groups. It's worth noting that while the generation of these projections is grounded in the generalized statistical …
Making Data Meaningful: Stakeholder Perceptions On Data Visualization And Data Management Practices Within A Multi-Tiered System Of Supports (Mtss), Domenick Saia
Dissertations
Data-driven decision-making and collaboration are core pillars of a multi-tiered system of supports (MTSS); however, timely and accessible data use, as well as data literacy and visualization literacy skills, are challenges school leaders and educators face related to implementing such frameworks. I hypothesized efficient data management systems and data visualization tools enable school teams to predict student learning outcomes, readily communicate, and better understand student data. The purpose of this study design was to highlight a need for more efficient data structures that allow school stakeholders to balance their roles within an MTSS framework more effectively. The context of this …
Analyzing Tortuosity In Patterns Formed By Colonies Of Embryonic Stem Cells Using Topological Data Analysis, Jackie Driscoll
Analyzing Tortuosity In Patterns Formed By Colonies Of Embryonic Stem Cells Using Topological Data Analysis, Jackie Driscoll
Master's Theses
Pluripotent stem cells have been observed to segregate into Turing-like patterns during the early stages of Dox-inducible hiPSC differentiation. In this thesis, we de- velop a tool to quantify the tortuosity in the patterns formed by colonies of pluripo- tent stem cells using methods from topological data analysis. We use clustering techniques and the mapper algorithm to create simplicial complexes representing samples of cells and detail a method of evaluating the tortuosity of these complexes. We use the resulting persistence landscapes and their associated norms to evaluate experimental data and simulated data from an agent based model. This thesis finds …