Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Smith College (49)
- Southern Methodist University (37)
- Kennesaw State University (27)
- Central Bank of Nigeria (23)
- Old Dominion University (23)
-
- City University of New York (CUNY) (19)
- University of Central Florida (16)
- Chapman University (13)
- West Virginia University (13)
- Department of Primary Industries and Regional Development, Western Australia (12)
- Illinois State University (12)
- California Polytechnic State University, San Luis Obispo (11)
- East Tennessee State University (11)
- LSU Health New Orleans (11)
- Claremont Colleges (10)
- Georgia Southern University (10)
- University of Arkansas, Fayetteville (10)
- Embry-Riddle Aeronautical University (9)
- University of Kentucky (9)
- Virginia Commonwealth University (9)
- Rochester Institute of Technology (7)
- Clemson University (6)
- Dartmouth College (6)
- Purdue University (6)
- Binghamton University (5)
- Murray State University (5)
- The University of Akron (5)
- University of Louisville (5)
- University of New Mexico (5)
- University of South Florida (5)
- Keyword
-
- Machine Learning (36)
- Machine learning (36)
- Statistics (27)
- Deep learning (15)
- Data Science (14)
-
- Data science (13)
- Classification (12)
- Time series (11)
- COVID-19 (10)
- Deep Learning (9)
- Artificial Intelligence (8)
- Forecasting (8)
- Regression (8)
- Prediction (7)
- Western Australia (7)
- Logistic regression (6)
- Neural Network (6)
- Clustering (5)
- Data analysis (5)
- Natural language processing (5)
- Sentiment analysis (5)
- Simulation (5)
- Time Series (5)
- Analysis (4)
- Analytics (4)
- Baseball (4)
- Bayesian (4)
- Bioinformatics (4)
- Biostatistics (4)
- CNN (4)
- Publication Year
- Publication
-
- Statistical and Data Sciences: Faculty Publications (45)
- SMU Data Science Review (27)
- CBN Journal of Applied Statistics (JAS) (20)
- Symposium of Student Scholars (20)
- Electronic Theses and Dissertations (18)
-
- Theses and Dissertations (17)
- Master's Theses (11)
- Annual Symposium on Biomathematics and Ecology Education and Research (10)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (10)
- Mathematics & Statistics Faculty Publications (10)
- School of Public Health Faculty Publications (10)
- College of Graduate Studies: Theses & Dissertations (9)
- Data Science and Data Mining (9)
- Dissertations, Theses, and Capstone Projects (9)
- Articles (7)
- Computational and Data Sciences (PhD) Dissertations (7)
- Dissertations (7)
- CMC Senior Theses (6)
- All Dissertations (5)
- Honors College Theses (5)
- Northeast Journal of Complex Systems (NEJCS) (5)
- Publications and Research (5)
- Statistical Science Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Dartmouth College Ph.D Dissertations (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Doctor of Data Science and Analytics Dissertations (4)
- Electronic Theses & Dissertations (2024 - present) (4)
- Fisheries Research Articles (4)
- Honors Projects (4)
- Publication Type
- File Type
Articles 151 - 180 of 550
Full-Text Articles in Data Science
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Calculation And Statistical Analysis Of Wins Above Replacement, Joshua Taylor
Departmental Honors & Graduate Capstone Projects
The Wins Above Replacement (WAR) statistic in Major League Baseball is a prominent metric used to estimate player value by quantifying all aspects of play in terms of wins added to a baseball team. We will use R to calculate WAR for all players from 1871 to 2012 and use data from those years to construct multivariate predictive models to attempt to estimate WAR for players from 2013 to 2024. We find strong correlations between predicted and actual WAR values for most models, with the exception of the polynomial predictive model for non-qualified pitchers.
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar
CBER Conference
Data is the fundamental building block for advancements in artificial intelligence (AI), general AI (GAI), machine learning (ML), and large language models (LLMs). This study emphasizes the critical need for robust data infrastructure, arguing that without it, countries cannot fully benefit from technological advancements in various economic sectors. Governments possess vast repositories of both structured and unstructured data across multiple domains such as the judiciary, parliaments, and civil bureaucracy. However, these potential goldmines remain untapped due to inadequate data management capabilities and a lack of appreciation for the necessity of high-quality data. The research identifies key issues in public data …
Customer Data And The Digital Age, Mahdi Ansari
Customer Data And The Digital Age, Mahdi Ansari
CBER Conference
Data is widely regarded as the most valuable resource in today’s economy, yet its value often eludes precise quantification. This paper examines customer data as an intangible capital asset and addresses the challenge of measuring its impact. A novel database was created by merging Compustat with online clickstream data capturing the activity of approximately 200 million users, providing proxies for data inflow based on visit metrics. The analysis documents that the distribution of firms’ customer data stocks follows a rightskewed log-normal pattern with a fat tail. Additionally, a positive relationship emerges between sales and data inflow, data stock, profit, and …
Statistical Analysis For Pre- And Post- Assessments Of Sdq And Idela Scores, Diego Murillo, Franceli L. Cibrian
Statistical Analysis For Pre- And Post- Assessments Of Sdq And Idela Scores, Diego Murillo, Franceli L. Cibrian
Student Scholar Symposium Abstracts and Posters
This research aimed to assess the potential of Mazi Umntanakho ("Know Your Child") in tracking developmental milestones in young children. Mazi is a WhatsApp-based conversational agent that assists South African home visitors in evaluating and monitoring children's socio-emotional skills using the Strengths and Difficulties Questionnaire (SDQ) and the International Development and Early Learning Assessment (IDELA). A field study was conducted in low-income South African communities, where 95 home visitors assessed 1,208 children. This detailed analysis of the data was collected during that deployment, focusing on investigating whether assessment scores improved over time and whether the length of time between assessments …
Optimization Of Markov Chain Modeling In Predicting College Student Retention, Kien Nguyen
Optimization Of Markov Chain Modeling In Predicting College Student Retention, Kien Nguyen
Journal of Global Education and Research
College student retention is one of the most important metrics in higher education. With institutions across the US facing decreasing enrollment, developing a reliable retention prediction method is crucial. In recent years, the use of the Markov chain model in forecasting student enrollment and progression has become more common, but there is little work on its application in student retention. One key factor in determining this model's effectiveness is what parameters should be used in the student population’s segmentation or grouping. This study presents a rigorous algorithm, coupled with a prediction model, capable of selecting parameters that provide the most …
Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni
Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni
Theses and Dissertations
This dissertation addresses critical challenges in neural network design by leveraging entropy-based techniques to improve model efficiency, interpretability, and bias reduction. Focusing on the unique demands of computer vision applications, particularly object detection and classification for real-time systems, this work introduces a series of innovative methods centered on information theory. At the core of these methods is the Probabilistic Explanations of Entropic Knowledge (PEEK) framework, a tool developed to analyze and visualize entropy distributions across feature maps. PEEK offers insights into information flow within neural networks, making it possible to pinpoint layers that contribute meaningfully to decision-making or identify those …
Status Quo Of Large-Scale Models, Risks And Challenges, And Recommended Countermeasures, Le Cheng, Yang Xiao
Status Quo Of Large-Scale Models, Risks And Challenges, And Recommended Countermeasures, Le Cheng, Yang Xiao
Bulletin of Chinese Academy of Sciences (Chinese Version)
Large-scale models (large models) are not only central to technological innovation, but also deeply entwined with national security, economic transformation, and social governance. This study examines the status quo of large-model development, identifies the key risks and challenges, and proposes response strategies, aiming to provide theoretical and policy insights for China’s navigations in global artificial intelligence (AI) competition and advances technological innovation. The research indicates that competition in the large-model market is fierce, while the industry is gradually consolidating. Competition in large models between China and the United States has escalated into a form of geopolitical contest. From a technical …
The Wallet And The Gut: Forecasting The 2024 Presidential Election With A State-By-State Adaptation Of The Time-For-Change Model, Simeon A. Betapudi, Hadassah Betapudi
The Wallet And The Gut: Forecasting The 2024 Presidential Election With A State-By-State Adaptation Of The Time-For-Change Model, Simeon A. Betapudi, Hadassah Betapudi
Science University Research Symposium (SURS)
This study adapts Abramowitz's Time-for-Change model to a state-level framework to forecast the 2024 U.S. presidential election. The Time-for-Change model’s focus on the popular vote has become less relevant in recent years, given the growing divergence between popular vote outcomes and electoral college results. Our model addresses these issues by adapting the original Time-for-Change predictors (presidential approval rating, GDP, and time in office) to the state level. Using data from five election cycles (2004–2020), we employ an Ordinary Least Squares (OLS) regression to predict incumbent two-party vote share. Unlike the original model, state-level GDP and incumbency duration were found to …
Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley
Making Plans Findable, Accessible, Interoperable, And Reusable With Data Infrastructure: A Search Engine For Constructing, Analyzing, And Visualizing Planning Documents, Lindsay Poirier, Dexter Antonio, Makenna Dettmann, Tiffany Eng, Jennifer Ganata, Sujoy Ghosh, Mirthala Lopez, Ranesh Karma, Asiya Natekal, Catherine Brinkley
Statistical and Data Sciences: Faculty Publications
Local land-use plans help guide future development, but it is often difficult to compare content across jurisdictions, making regional coordination and plan evaluation challenging. This research reviews federal, state, and local data infrastructure guidance for land-use plans and compares such guidance to compliance with a California use-case. Findings indicate a number of obstacles to fostering data sharing and comparative analysis of plans: there is currently no central repository of land-use plans; plans are not uniform in format and are often out of date; many plans are not machine-readable thereby inhibiting text extraction, and planning language varies so greatly that there …
Cluster Effect For Snp-Snp Interaction Pairs For Predicting Complex Traits, Hui Yi Lin, Harun Mazumder, Indrani Sarkar, Po Yu Huang, Rosalind A. Eeles, Zsofia Kote-Jarai, Kenneth R. Muir, Johanna Schleutker, Nora Pashayan, Jyotsna Batra, David E. Neal, Sune F. Nielsen, Børge G. Nordestgaard, Henrik Grönberg, Fredrik Wiklund, Robert J. Macinnis, Christopher A. Haiman, Ruth C. Travis, Janet L. Stanford, Adam S. Kibel, Cezary Cybulski, Kay Tee Khaw, Christiane Maier, Stephen N. Thibodeau, Manuel R. Teixeira, Lisa Cannon-Albright, Hermann Brenner, Radka Kaneva, Hardev Pandha, Et Al
Cluster Effect For Snp-Snp Interaction Pairs For Predicting Complex Traits, Hui Yi Lin, Harun Mazumder, Indrani Sarkar, Po Yu Huang, Rosalind A. Eeles, Zsofia Kote-Jarai, Kenneth R. Muir, Johanna Schleutker, Nora Pashayan, Jyotsna Batra, David E. Neal, Sune F. Nielsen, Børge G. Nordestgaard, Henrik Grönberg, Fredrik Wiklund, Robert J. Macinnis, Christopher A. Haiman, Ruth C. Travis, Janet L. Stanford, Adam S. Kibel, Cezary Cybulski, Kay Tee Khaw, Christiane Maier, Stephen N. Thibodeau, Manuel R. Teixeira, Lisa Cannon-Albright, Hermann Brenner, Radka Kaneva, Hardev Pandha, Et Al
School of Public Health Faculty Publications
Single nucleotide polymorphism (SNP) interactions are the key to improving polygenic risk scores. Previous studies reported several significant SNP-SNP interaction pairs that shared a common SNP to form a cluster, but some identified pairs might be false positives. This study aims to identify factors associated with the cluster effect of false positivity and develop strategies to enhance the accuracy of SNP-SNP interactions. The results showed the cluster effect is a major cause of false-positive findings of SNP-SNP interactions. This cluster effect is due to high correlations between a causal pair and null pairs in a cluster. The clusters with a …
Disparities And Protective Factors In Pandemic-Related Mental Health Outcomes: A Louisiana-Based Study, Ariane L. Rung, Evrim Oral, Tyler Prusisz, Edward S. Peters
Disparities And Protective Factors In Pandemic-Related Mental Health Outcomes: A Louisiana-Based Study, Ariane L. Rung, Evrim Oral, Tyler Prusisz, Edward S. Peters
School of Public Health Faculty Publications
Introduction: The COVID-19 pandemic has had a wide-ranging impact on mental health. Diverse populations experienced the pandemic differently, highlighting pre-existing inequalities and creating new challenges in recovery. Understanding the effects across diverse populations and identifying protective factors is crucial for guiding future pandemic preparedness. The objectives of this study were to (1) describe the specific COVID-19-related impacts associated with general well-being, (2) identify protective factors associated with better mental health outcomes, and (3) assess racial disparities in pandemic impact and protective factors. Methods: A cross-sectional survey of Louisiana residents was conducted in summer 2020, yielding a sample of 986 Black …
Uncertainty Quantification In Machine Learning Models Via Gaussian Process Regression: A Comparative Study, Ayorinde E. Olatunde, Weiqi Yue, Pawan K. Tripathi, Roger H. French, Anirban Mondal
Uncertainty Quantification In Machine Learning Models Via Gaussian Process Regression: A Comparative Study, Ayorinde E. Olatunde, Weiqi Yue, Pawan K. Tripathi, Roger H. French, Anirban Mondal
Faculty Scholarship
As the use of Machine learning models in science and engineering continues to increase, there is an increasing need for quantifying the uncertainties inherent in the predictions of these models. The more complex a model is, the more the uncertainties in its predictions increase. Amongst the plethora of methodologies used in quantifying uncertainties lies Gaussian Process Regression (GPR). GPR surmounts some of the popular shortfalls of other state-of-the-art methodologies. Although GPR has some quick wins in its application for uncertainty quantification, it is plagued with some shortfalls, such as scalability issues when the feature space increases as well as an …
Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn
Forecasting Commercial Vehicle Miles Traveled (Vmt) In Urban California Areas, Steve Chung, Jaymin Kwon, Yushin Ahn
Mineta Transportation Institute
This study investigates commercial truck vehicle miles traveled (VMT) across six diverse California counties from 2000 to 2020. The counties—Imperial, Los Angeles, Riverside, San Bernardino, San Diego, and San Francisco—represent a broad spectrum of California’s demographics, economies, and landscapes. Using a rich dataset spanning demographics, economics, and pollution variables, we aim to understand the factors influencing commercial VMT. We first visually represent the geographic distribution of the counties, highlighting their unique characteristics. Linear regression models, particularly the least absolute shrinkage and selection operator (LASSO) and elastic net regressions are employed to identify key predictors of total commercial VMT. LASSO regression …
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
All Theses
High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …
High Fat Diet & Social Isolation: Interactive Effects On Pain, Cognition, & Neuroinflammation, Ian M. Campuzano
High Fat Diet & Social Isolation: Interactive Effects On Pain, Cognition, & Neuroinflammation, Ian M. Campuzano
Research Psychology Theses
Prior research has established a role for both social isolation and exposure to high fat Western diets in altering a range of behaviors from reduced memory performance to increased depression-like behaviors. The present study scrutinizes the interplay among these variables during the peri-adolescent developmental phase, utilizing Long-Evans rats as the experimental model. Our overarching hypothesis is that rats exposed to either social isolation, a high-fat diet, or both will result in heightened pain sensitivity, diminished cognitive flexibility, and increased neuroinflammatory responses within brain regions implicated in sociability, cognition, memory, and pain processing. Behavioral flexibility will be assessed using a maze-based …
Exploring The Diagnostic Potential Of Radiomics-Based Pet Image Analysis For T-Stage Tumor Diagnosis, Victor Aderanti
Exploring The Diagnostic Potential Of Radiomics-Based Pet Image Analysis For T-Stage Tumor Diagnosis, Victor Aderanti
Electronic Theses and Dissertations
Cancer is a leading cause of death globally, and early detection is crucial for better
outcomes. This research aims to improve Region Of Interest (ROI) segmentation
and feature extraction in medical image analysis using Radiomics techniques
with 3D Slicer, Pyradiomics, and Python. Dimension reduction methods, including
PCA, K-means, t-SNE, ISOMAP, and Hierarchical Clustering, were applied to highdimensional features to enhance interpretability and efficiency. The study assessed the ability of the reduced feature set to predict T-staging, an essential component of the TNM system for cancer diagnosis. Multinomial logistic regression models were developed and evaluated using MSE, AIC, BIC, and Deviance …
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
All Dissertations
The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …
Gender-Specific Mental Health Outcomes In Central America: A Natural Experiment, Thea Nagasuru
Gender-Specific Mental Health Outcomes In Central America: A Natural Experiment, Thea Nagasuru
Computer Science Summer Fellows
While COVID lockdown measures have had varying effects on the mental health of different demographics, several bodies of research have noted their disparate effect on women. Why is women's mental health more negatively impacted by lockdown measures, and how much more are they impacted than men? How can we predict and mitigate these negative effects on women? This paper aims to contribute to answering those questions by comparing COVID stringency measures and their effect on the gap in depression rates between men and women in two neighboring countries: Nicaragua and Honduras.
Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier
Enacting Data Context: Fixing Meaning In Transparency Data Initiatives, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
This article documents the “context cultures” underpinning efforts to develop regulations for collecting and reporting data in a United States public database known as Open Payments. Open Payments is a dataset published annually by the US Center for Medicare and Medicaid Services that documents the transfers of value from pharmaceutical and medical device manufacturers to physicians, prescribing non-physicians, and teaching hospitals. In the article, I show context became a manifold concern as differentially-situated actors engaged in modes of public advocacy and social action around not only what data meant, but also what it meant to make data meaningful. I show …
Extreme Value Statistics Analysis Of Process Defects In Additive Manufacturing Materials, Ayorinde E. Olatunde, Kristen Hernandez, Austin Ngo, Arafath Nihar, Thomas G. Ciardi, Rachel Yamamoto, Pawan K. Tripathi, Roger H. French, John J. Lewandowski, Anirban Mondal
Extreme Value Statistics Analysis Of Process Defects In Additive Manufacturing Materials, Ayorinde E. Olatunde, Kristen Hernandez, Austin Ngo, Arafath Nihar, Thomas G. Ciardi, Rachel Yamamoto, Pawan K. Tripathi, Roger H. French, John J. Lewandowski, Anirban Mondal
Faculty Scholarship
Fatigue and fracture studies focused on process defects that occur in Additive Manufacturing (AM) materials have shown that defect populations possess features which are better measured with extreme value statistics (EVS). In AM alloys, defect occurrences increase with material volume. This situation facilitates the need to model process defects in the path of fatigue crack growth with suitable statistical tools, such as EVS, which is more cost-effective when compared to destructive experiments. The application of EVS on defect space features helps determine the difference in defects present on fracture surfaces. As the fatigue quality of any material depends on its …
Examining The Interaction Between Calcium Supplement Use, Demographics, And Lifestyle Factors On Bone Health In Women, Vix Talbot
University Honors Theses
Osteoporosis is a condition which poses a significant health threat, particularly among women during the menopause transition, where accelerated bone loss increases fracture risk. Calcium supplementation has been shown to be an important intervention to mitigate bone mineral density (BMD) decline during this and other periods of life. However, the efficacy of calcium supplementation is influenced by various individual factors, including demographics and lifestyle habits. This study investigates the interaction between calcium supplement use, and several interaction terms on bone health in women. Multiple linear regression analysis is employed to assess the impact of these factors on BMD. Data from …
Accessible Real-Time Eye-Gaze Tracking For Neurocognitive Health Assessments, A Multimodal Web-Based Approach, Daniel C. Tisdale
Accessible Real-Time Eye-Gaze Tracking For Neurocognitive Health Assessments, A Multimodal Web-Based Approach, Daniel C. Tisdale
Master's Theses
We introduce a novel integration of real-time, predictive eye-gaze tracking models into a multimodal dialogue system tailored for remote health assessments. This system is designed to be highly accessible requiring only a conventional webcam for video input along with minimal cursor interaction and utilizes engaging gaze-based tasks that can be performed directly in a web browser. We have crafted dynamic subsystems that capture high-quality data efficiently and maintain quality through instances of user attrition and incomplete calls. Additionally, these subsystems are designed with the foresight to allow for future re-analysis using improved predictive models, as well as enable the creation …
Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov
Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov
CBN Journal of Applied Statistics (JAS)
This paper investigates the time it would take for the FTSE-100 index to reach its post-COVID-19 peak. The paper utilises an exponential generalised autoregressive conditional heteroscedasticity (EGARCH) model that accounts for leverage effect and asymmetries. The preferred models amongst competing variants was the Autoregressive Moving Average (ARMA)-EGARCH(2,1) specification and was used to predict daily FTSE-100 data from 5th January 2000 to 21st June 2024. The empirical exercise showed that the COVID-19-induced financial crisis negatively affected the United Kingdom’s stock market performance. The results show that the FTSE100 index could reach its post-pandemic peak around 27th August, 2024 (two months after …
A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte
A Symbolic Approach To Nonlinear Time Series Analysis, Ranjan Karki, Nibhrat Lohia, Michael B. Schulte
SMU Data Science Review
Current nonlinear time series methods such as neural networks forecast well. However, they act as a black box and are difficult to interpret, leaving the researchers and the audience with little insight into why the forecasts are the way they are. There is a need for a method that forecasts accurately while also being easy to interpret. This paper aims to develop a method to build an interpretable model for univariate and multivariate nonlinear time series data using wavelets and symbolic regression. The final method relies on multilayer perceptron (MLP) neural networks as a form of dimensionality reduction and the …
Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler Ii, Alexandria Neff, Adam R. Ruthford
Reevaluating Texas Energy Market Forecasts In The Wake Of Recent Extreme Weather Events, Robert A. Derner, Richard W. Butler Ii, Alexandria Neff, Adam R. Ruthford
SMU Data Science Review
This paper provides updated forecasts of energy demand in Texas and recognizes the impact of sustainable energy. It is important that the forecasts of the adoption of sustainable energy are reexamined after Winter Storm Uri crippled the Texas power grid and left many without power. This storm highlighted the issues the Texas power grid had and has continued to struggle with in supplying the state with energy. This paper will offer an overview of the relevant literature on the adoption of sustainable energy and relevant events that have occurred in the state of Texas that will give the reader the …
Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn
Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn
SMU Data Science Review
As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …
Context Aware Music Recommendation And Playlist Generation, Elias Mann
Context Aware Music Recommendation And Playlist Generation, Elias Mann
SMU Journal of Undergraduate Research
There are many reasons people listen to music, and the type of music is largely determined by what the listener may be doing while they listen. For example, one may listen to one type of music while commuting, another while exercising, and yet another while relaxing. Without access to the physiological state of the user, current music recommendation methods rely on collaborative filtering - recommending music based on what other similar users listen to - and content based filtering - recommending songs based on their similarities to songs the user already prefers. With the rise in popularity of smart devices …
Capturing Higher-Order Relationships Through Information Decomposition, Aobo Lyu
Capturing Higher-Order Relationships Through Information Decomposition, Aobo Lyu
McKelvey School of Engineering Graduate Student Theses & Dissertations
Mutual information between two random variables is a well-studied notion, whose understanding is fairly complete. Mutual information between one random variable and a pair of other random variables, however, is a far more involved notion. Specifically, Shannon's mutual information does not capture fine-grained interactions between those three variables, resulting in limited insights in complex systems. To capture these fine-grained higher-order interactions among variables, Williams and Beer proposed a framework called Partial Information Decomposition (PID) to decompose this mutual information to information atoms, called unique, redundant, and synergistic, and proposed several operational axioms that these atoms must satisfy. This conceptual …
A Spatial Decision Support System For Rent Estimation Of Retail Spaces In Manhattan Using Geographically Weighted Regression And Spatial Regression, Andie M. Migden Miller
A Spatial Decision Support System For Rent Estimation Of Retail Spaces In Manhattan Using Geographically Weighted Regression And Spatial Regression, Andie M. Migden Miller
Theses and Dissertations
This report outlines an automated, three-phase Spatial Decision Support System that creates models to estimate rent of retail spaces across Manhattan. First, enrich data with predictors. Second, optimize spatially aware neighborhood-level models by combining GWR, spatial regression, and non-spatial regression. Finally, visualize results in an Esri-based WebApp.
A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang
A Novel Correction For The Multivariate Ljung-Box Test, Minhao Huang
Computational and Data Sciences (PhD) Dissertations
This research introduces an analytical improvement to the Multivariate Ljung-Box test that addresses significant deviations of the original test from the nominal Type I error rates under almost all scenarios. Prior attempts to mitigate this issue have been directed at modification of the test statistics or correction of the test distribution to achieve precise results in finite samples. In previous studies, focused on designing corrections to the univariate Ljung-Box, a method that specifically adjusts the test rejection region has been the most successful of attaining the best Type I error rates. We adopt the same approach for the more complex, …