Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Smith College (49)
- Southern Methodist University (37)
- Kennesaw State University (27)
- Central Bank of Nigeria (23)
- Old Dominion University (23)
-
- City University of New York (CUNY) (19)
- University of Central Florida (16)
- Chapman University (13)
- West Virginia University (13)
- Department of Primary Industries and Regional Development, Western Australia (12)
- Illinois State University (12)
- California Polytechnic State University, San Luis Obispo (11)
- East Tennessee State University (11)
- LSU Health New Orleans (11)
- Claremont Colleges (10)
- Georgia Southern University (10)
- University of Arkansas, Fayetteville (10)
- Embry-Riddle Aeronautical University (9)
- University of Kentucky (9)
- Virginia Commonwealth University (9)
- Rochester Institute of Technology (7)
- Clemson University (6)
- Dartmouth College (6)
- Purdue University (6)
- Binghamton University (5)
- Murray State University (5)
- The University of Akron (5)
- University of Louisville (5)
- University of New Mexico (5)
- University of South Florida (5)
- Keyword
-
- Machine Learning (36)
- Machine learning (36)
- Statistics (27)
- Deep learning (15)
- Data Science (14)
-
- Data science (13)
- Classification (12)
- Time series (11)
- COVID-19 (10)
- Deep Learning (9)
- Artificial Intelligence (8)
- Forecasting (8)
- Regression (8)
- Prediction (7)
- Western Australia (7)
- Logistic regression (6)
- Neural Network (6)
- Clustering (5)
- Data analysis (5)
- Natural language processing (5)
- Sentiment analysis (5)
- Simulation (5)
- Time Series (5)
- Analysis (4)
- Analytics (4)
- Baseball (4)
- Bayesian (4)
- Bioinformatics (4)
- Biostatistics (4)
- CNN (4)
- Publication Year
- Publication
-
- Statistical and Data Sciences: Faculty Publications (45)
- SMU Data Science Review (27)
- CBN Journal of Applied Statistics (JAS) (20)
- Symposium of Student Scholars (20)
- Electronic Theses and Dissertations (18)
-
- Theses and Dissertations (17)
- Master's Theses (11)
- Annual Symposium on Biomathematics and Ecology Education and Research (10)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (10)
- Mathematics & Statistics Faculty Publications (10)
- School of Public Health Faculty Publications (10)
- College of Graduate Studies: Theses & Dissertations (9)
- Data Science and Data Mining (9)
- Dissertations, Theses, and Capstone Projects (9)
- Articles (7)
- Computational and Data Sciences (PhD) Dissertations (7)
- Dissertations (7)
- CMC Senior Theses (6)
- All Dissertations (5)
- Honors College Theses (5)
- Northeast Journal of Complex Systems (NEJCS) (5)
- Publications and Research (5)
- Statistical Science Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Dartmouth College Ph.D Dissertations (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Doctor of Data Science and Analytics Dissertations (4)
- Electronic Theses & Dissertations (2024 - present) (4)
- Fisheries Research Articles (4)
- Honors Projects (4)
- Publication Type
- File Type
Articles 361 - 390 of 550
Full-Text Articles in Data Science
An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel
An Educator’S Perspective Of The Tidyverse, Mine Çetinkaya-Rundel, Johanna Hardin, Benjamin Baumer, Amelia Mcnamara, Nicholas J. Horton, Colin W. Rundel
Statistical and Data Sciences: Faculty Publications
Computing makes up a large and growing component of data science and statistics courses. Many of those courses, especially when taught by faculty who are statisticians by training, teach R as the programming language. A number of instructors have opted to build much of their teaching around use of the tidyverse. The tidyverse, in the words of its developers, “is a collection of R packages that share a high-level design philosophy and low-level grammar and data structures, so that learning one package makes it easier to learn the next” (Wickham et al. 2019). These shared principles have led to the …
Intra-Hour Solar Forecasting Using Cloud Dynamics Features Extracted From Ground-Based Infrared Sky Images, Guillermo Terrén-Serrano
Intra-Hour Solar Forecasting Using Cloud Dynamics Features Extracted From Ground-Based Infrared Sky Images, Guillermo Terrén-Serrano
Electrical and Computer Engineering ETDs
Due to the increasing use of photovoltaic systems, power grids are vulnerable to the projection of shadows from moving clouds. An intra-hour solar forecast provides power grids with the capability of automatically controlling the dispatch of energy, reducing the additional cost for a guaranteed, reliable supply of energy (i.e., energy storage). This dissertation introduces a novel sky imager consisting of a long-wave radiometric infrared camera and a visible light camera with a fisheye lens. The imager is mounted on a solar tracker to maintain the Sun in the center of the images throughout the day, reducing the scattering effect produced …
A New Application Of The Central Limit Theorem, Kenneth Winters
A New Application Of The Central Limit Theorem, Kenneth Winters
Selected Honors Theses
This paper discusses the Central Limit Theorem (CLT) and its applications. The paper gives an introduction to what the CLT is and how it can be applied to real life. Additionally, the paper gives a conceptual understanding of the theorem through various examples and visuals. The paper discusses the applications of the CLT in fields such as computer science, psychology, and political science. The author then suggests a new mathematical theorem as an application of the CLT and provides a proof of the theorem. The new theorem relates to expected value and probabilities of random variables and provides a link …
Split Classification Model For Complex Clustered Data, Katherine Gerot
Split Classification Model For Complex Clustered Data, Katherine Gerot
Honors Program: Senior Projects (Public)
Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.
Learning Latent Causal Dynamics, Weiran Yao, Guangyi Chen, Kun Zhang
Learning Latent Causal Dynamics, Weiran Yao, Guangyi Chen, Kun Zhang
Machine Learning Faculty Publications
One critical challenge of time-series modeling is how to learn and quickly correct the model under unknown distribution shifts. In this work, we propose a principled framework, called LiLY, to first recover time-delayed latent causal variables and identify their relations from measured temporal data under different distribution shifts. The correction step is then formulated as learning the low-dimensional change factors with a few samples from the new environment, leveraging the identified causal structure. Specifically, the framework factorizes unknown distribution shifts into transition distribution changes caused by fixed dynamics and time-varying latent causal relations, and by global changes in observation. We …
Session 5: Equipment Finance Credit Risk Modeling - A Case Study In Creative Model Development & Nimble Data Engineering, Edward Krueger, Landon Thompson, Josh Moore
Session 5: Equipment Finance Credit Risk Modeling - A Case Study In Creative Model Development & Nimble Data Engineering, Edward Krueger, Landon Thompson, Josh Moore
SDSU Data Science Symposium
This presentation will focus first on providing an overview of Channel and the Risk Analytics team that performed this case study. Given that context, we’ll then dive into our approach for building the modeling development data set, techniques and tools used to develop and implement the model into a production environment, and some of the challenges faced upon launch. Then, the presentation will pivot to the data engineering pipeline. During this portion, we will explore the application process and what happens to the data we collect. This will include how we extract & store the data along with how it …
Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch
Mental Health In The Uk Biobank: A Roadmap To Self-Report Measures And Neuroimaging Correlates, Rosie K. Dutt, Kayla Hannon, Ty O. Easley, Joseph C. Griffis, Wei Zhang, Janine D. Bijsterbosch
Statistical and Data Sciences: Faculty Publications
The UK Biobank (UKB) is a highly promising dataset for brain biomarker research into population mental health due to its unprecedented sample size and extensive phenotypic, imaging, and biological measurements. In this study, we aimed to provide a shared foundation for UKB neuroimaging research into mental health with a focus on anxiety and depression. We compared UKB self-report measures and revealed important timing effects between scan acquisition and separate online acquisition of some mental health measures. To overcome these timing effects, we introduced and validated the Recent Depressive Symptoms (RDS-4) score which we recommend for state-dependent and longitudinal research in …
Liquidity Commonality With Factor Models, Ernesto Garcia Iii
Liquidity Commonality With Factor Models, Ernesto Garcia Iii
Dissertations, Theses, and Capstone Projects
Market microstructure research has recently devoted attention to a phenomenon called commonality in liquidity. In this dissertation, I will analyze commonality in liquidity using a novel factor model approach and a generalized definition of commonality in liquidity. This analysis will show that commonality in liquidity is rarely a marketwide phenomenon and is mostly restricted to stocks with a large market capitalization. Additionally, commonality in liquidity is a very recent phenomenon whose appearance coincides with a rise in passive investing after the Dotcom Bubble burst and, more so, after the 2008 Financial Crisis. I will present evidence that suggests commonality in …
The Data Analytics And The Science Revolution, Leila Halawi, Amal Clarke, Kelly George
The Data Analytics And The Science Revolution, Leila Halawi, Amal Clarke, Kelly George
Publications
This text highlights the difference between analytics and data science, using predictive analytic techniques to analyze different historical data, including aviation data and concrete data, interpreting the predictive models, and highlighting the steps to deploy the models and the steps ahead. The book combines the conceptual perspective and a hands-on approach to predictive analytics using SAS VIYA, an analytic and data management platform. The authors use SAS VIYA to focus on analytics to solve problems, highlight how analytics is applied in the airline and business environment, and compare several different modeling techniques. They decipher complex algorithms to demonstrate how they …
Transition Metal Phosphides For High Performance Electrochemical Energy Storage Devices, Amina Saleh
Transition Metal Phosphides For High Performance Electrochemical Energy Storage Devices, Amina Saleh
Theses and Dissertations
Electrochemical energy storage technologies are nowadays playing a leading role in the global effort to address the energy challenges. A lot of attention has been devoted to designing hybrid devices known as supercapatteries which combine the merits of supercapacitors (high power density) and rechargeable batteries (high energy density). Transition metal phosphides (TMP) are a rising star for supercapattery anode materials thanks to their high conductivity, metalloid characteristics, and kinetic favorability for fast electron transport. Herein, new TMP-based materials were synthesized for use as supercapattery positive electrodes, via a multifaceted approach to yield devices enjoying concurrently high power and energy densities. …
Author’S Reflections On Making Sense Of Numbers: Quantitative Reasoning For Social Research, Jane E. Miller
Author’S Reflections On Making Sense Of Numbers: Quantitative Reasoning For Social Research, Jane E. Miller
Numeracy
Miller, Jane E. 2021. Making Sense of Numbers: Quantitative Reasoning for Social Research. (Los Angeles: SAGE Publications) 608 pp. ISBN 978-1544355597.
This article introduces and provides an excerpt from Making Sense of Numbers: Quantitative Reasoning for Social Research, published by Sage. The book explains and illustrates how making sense of numbers involves integrating concepts and skills from mathematics, statistics, study design, and communications, along with information about the specific topic and context under study. It teaches how to avoid making common errors of logic, calculation, and interpretation by introducing a systematic approach and a healthy dose of skepticism …
A Predictive Model To Predict Cyberattack Using Self-Normalizing Neural Networks, Oluwapelumi Eniodunmo
A Predictive Model To Predict Cyberattack Using Self-Normalizing Neural Networks, Oluwapelumi Eniodunmo
Theses, Dissertations and Capstones
Cyberattack is a never-ending war that has greatly threatened secured information systems. The development of automated and intelligent systems provides more computing power to hackers to steal information, destroy data or system resources, and has raised global security issues. Statistical and Data mining tools have received continuous research and improvements. These tools have been adopted to create sophisticated intrusion detection systems that help information systems mitigate and defend against cyberattacks. However, the advancement in technology and accessibility of information makes more identifiable elements that can be used to gain unauthorized access to systems and resources. Data mining and classification tools …
Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward
Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward
College of Graduate Studies: Theses & Dissertations
In non-inferiority testing, the decision of whether a proposed treatment is non-inferior to a reference treatment depends on model assumptions and choices of acceptable tolerance limits. Here, we consider a method that employs kernels to estimate the probability density functions of both the experimental and reference populations from two independent samples. Based on these densities, we introduce a quantity called the overlap coefficient or overlap measure. A bootstrap technique is helpful in exploring the distribution and variance empirically. We derive the distribution of this measure and define a hypothesis test that can be applied to the non-inferiority setting under some …
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Honors Theses and Capstones
COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …
Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su
Theses and Dissertations--Statistics
When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …
Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton
Addressing Ascertainment Bias In The Study Of Cardiovascular Disease Burden In Opioid Use Disorders - Application Of Natural Language Processing Of Electronic Health Records, Jade Huang Singleton
Theses and Dissertations--Epidemiology and Biostatistics
In the United States, the prevalence of long-term exposure to opioid drugs, for both medically and nonmedically indicated purposes, has increased considerably since the mid-1990’s. Concerns have emerged about the potential health effects of opioid use. There is also growing interest in other possible connections with opioid use including cardiovascular disease. Electronic health records (EHR) contain information about patient care in the form of structured codes and unstructured notes. Natural language processing (NLP) provides a tool for processing unstructured textual data in EHR clinical notes and extracts useful information for research with structured formats. The purpose of this dissertation was …
Realtime Event Detection In Sports Sensor Data With Machine Learning, Mallory Cashman
Realtime Event Detection In Sports Sensor Data With Machine Learning, Mallory Cashman
Honors Theses and Capstones
Machine learning models can be trained to classify time series based sports motion data, without reliance on assumptions about the capabilities of the users or sensors. This can be applied to predict the count of occurrences of an event in a time period. The experiment for this research uses lacrosse data, collected in partnership with SPAITR - a UNH undergraduate startup developing motion tracking devices for lacrosse. Decision Tree and Support Vector Machine (SVM) models are trained and perform with high success rates. These models improve upon previous work in human motion event detection and can be used a reference …
Data, Knowledge Practices, And Naturecultural Worlds: Vehicle Emissions In The Anthropocene, Lindsay Poirier
Data, Knowledge Practices, And Naturecultural Worlds: Vehicle Emissions In The Anthropocene, Lindsay Poirier
Statistical and Data Sciences: Faculty Books
This chapter details the various techno-cultural assemblages giving rise to data collected to model and measure anthropogenic worlds, arguing that data-based technologies both represent and co-produce the Anthropocene. It begins with a review of scholarship emerging at the intersection of science and technology studies and information studies that advances understanding of data infrastructure and knowledge practices, and their role within the anthropogenic assemblages that shape history. Drawing on a case study describing how vehicle emissions are measured and regulated in the US, I examine the materialities and mutability of technologies designed to produce data about air quality, along with the …
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Comparison Of Caregiver- And Child-Reported Quality Of Life In Children With Sleep-Disordered Breathing, Phoebe Kuo Yu, Kaitlyn Cook, Jiayan Liu, Raouf S. Amin, Craig Derkay, Lisa M. Elden, Susan L. Garetz, Alisha S. George, Sally Ibrahim, Stacey L. Ishman, Erin M. Kirkham, S. Kamal Naqvi, Jerilynn Radcliffe, Kristie R. Ross, Gopi B. Shah, Ignacio E. Tapia, H. Gerry Taylor, David A. Zopf, Susan Redline, Cristina M. Baldassari
Statistical and Data Sciences: Faculty Publications
Objective. Caregivers frequently report poor quality of life(QOL) in children with sleep-disordered breathing (SDB).Our objective is to assess the correlation between care-giver- and child-reported QOL in children with mild SDBand identify factors associated with differences between caregiver and child report.
Study Design. Analysis of baseline data from a multi-institutional randomized trialSetting. Pediatric Adenotonsillectomy Trial for Snoring, where children with mild SDB (obstructive apnea-hypopnea index\3) were randomized to observation or adenotonsillectomy.
Methods. The Pediatric Quality of Life Inventory (Peds QL)assessed baseline global QOL in participating children 5 to12 years old and their caregivers. Caregiver and child scores were compared. Multivariable regression …
Graph Neural Networks For Improved Interpretability And Efficiency, Patrick Pho
Graph Neural Networks For Improved Interpretability And Efficiency, Patrick Pho
Electronic Theses and Dissertations, 2020-2023
Attributed graph is a powerful tool to model real-life systems which exist in many domains such as social science, biology, e-commerce, etc. The behaviors of those systems are mostly defined by or dependent on their corresponding network structures. Graph analysis has become an important line of research due to the rapid integration of such systems into every aspect of human life and the profound impact they have on human behaviors. Graph structured data contains a rich amount of information from the network connectivity and the supplementary input features of nodes. Machine learning algorithms or traditional network science tools have limitation …
Change Point Detection For Streaming Data Using Support Vector Methods, Charles Harrison
Change Point Detection For Streaming Data Using Support Vector Methods, Charles Harrison
Electronic Theses and Dissertations, 2020-2023
Sequential multiple change point detection concerns the identification of multiple points in time where the systematic behavior of a statistical process changes. A special case of this problem, called online anomaly detection, occurs when the goal is to detect the first change and then signal an alert to an analyst for further investigation. This dissertation concerns the use of methods based on kernel functions and support vectors to detect changes. A variety of support vector-based methods are considered, but the primary focus concerns Least Squares Support Vector Data Description (LS-SVDD). LS-SVDD constructs a hypersphere in a kernel space to bound …
Applying Machine Learning Algorithms For Face Mask Detections, Mackenzie Frato
Applying Machine Learning Algorithms For Face Mask Detections, Mackenzie Frato
Williams Honors College, Honors Research Projects
Goal: Apply multiple machine learning techniques to Face Mask images to detect if a student is wear a Face Mask and/or wearing it incorrectly or not at all. Methodology: Use 2-3 different machine learning techniques to develop this program. Will choose these techniques as I research over the semester. The best technique will be the final one used, but many will be explored. Validation techniques will be used to see which is the best technique. Timeline: Choose Dataset - October 1st, Choose techniques - October 31st, Research techniques/validation - November 31st, Begin writing code - December 13th, Finish code - …
Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling
Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling
Williams Honors College, Honors Research Projects
This study uses various statistical analyses to evaluate the justification of rule changes for Major League Baseball that were implemented within the Minor Leagues during the 2021 minor league season. The primary focus of the study is predicting how some of these Minor League rule changes could affect the stolen base success rate and the number of attempts per game within the Major Leagues. A survey was conducted to evaluate how fans feel about stolen bases within the current game and if rules should be altered to increase the number of stolen bases that occur. Additionally, recorded Major and Minor …
Exploring Cyberterrorism, Topic Models And Social Networks Of Jihadists Dark Web Forums: A Computational Social Science Approach, Vivian Fiona Guetler
Exploring Cyberterrorism, Topic Models And Social Networks Of Jihadists Dark Web Forums: A Computational Social Science Approach, Vivian Fiona Guetler
Graduate Theses, Dissertations, and Problem Reports (ETD)
This three-article dissertation focuses on cyber-related topics on terrorist groups, specifically Jihadists’ use of technology, the application of natural language processing, and social networks in analyzing text data derived from terrorists' Dark Web forums. The first article explores cybercrime and cyberterrorism. As technology progresses, it facilitates new forms of behavior, including tech-related crimes known as cybercrime and cyberterrorism. In this article, I provide an analysis of the problems of cybercrime and cyberterrorism within the field of criminology by reviewing existing literature focusing on (a) the issues in defining terrorism, cybercrime, and cyberterrorism, (b) ways that cybercriminals commit a crime in …
A Monte Carlo Simulation Of Rat Choice Behavior With Interdependent Outcomes, Michelle A. Frankot
A Monte Carlo Simulation Of Rat Choice Behavior With Interdependent Outcomes, Michelle A. Frankot
Graduate Theses, Dissertations, and Problem Reports (ETD)
Preclinical behavioral neuroscience often uses choice paradigms to capture psychiatric symptoms. In particular, the subfield of operant research produces nested datasets with many discrete choices in a session. The standard analytic practice is to aggregate choice into a continuous variable and analyze using ANOVA or linear regression. However, choice data often have multiple interdependent outcomes of interest, violating an assumption of general linear models. The aim of the current study was to quantify the accuracy of linear mixed-effects regression (LMER) for analyzing data from a 4-choice operant task called the Rodent Gambling Task (RGT), which measures decision-making in the context …
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Theses and Dissertations
Primary Care is on the frontlines of healthcare, thus they see the most diverse set of patients. In order to achieve high functioning primary care, a practice must establish empanelment, the pairing of patients to providers. Enumeration of empanelment, or estimating panel sizes, helps ensure that the demands of the patients demand the supply of providers and optimize the balance of primary care resources to improve quality of care. Further we can adjust panel sizes by using patient-level data on healthcare utilization and complexity extracted from the electronic medial record to determine the amount of care or burden of work …
Aspect-Based Sentiment Analysis Of Movie Reviews, Samuel Onalaja, Eric Romero, Bosang Yun
Aspect-Based Sentiment Analysis Of Movie Reviews, Samuel Onalaja, Eric Romero, Bosang Yun
SMU Data Science Review
This study investigates a comparison of classification models used to determine aspect based separated text sentiment and predict binary sentiments of movie reviews with genre and aspect specific driving factors. To gain a broader classification analysis, five machine and deep learning algorithms were compared: Logistic Regression (LR), Naive Bayes (NB), Support Vector Machine (SVM), and Recurrent Neural Network Long-Short-Term Memory (RNN LSTM). The various movie aspects that are utilized to separate the sentences are determined through aggregating aspect words from lexicon-base, supervised and unsupervised learning. The driving factors are randomly assigned to various movie aspects and their impact tied to …
Predicting Power Using Time Series Analysis Of Power Generation And Consumption In Texas, Joshua Eysenbach, Bodie Franklin, Andrew J. Larsen, Joel Lindsey
Predicting Power Using Time Series Analysis Of Power Generation And Consumption In Texas, Joshua Eysenbach, Bodie Franklin, Andrew J. Larsen, Joel Lindsey
SMU Data Science Review
Due to the recent power events in Texas, power forecasting has been brought national attention. Accurate demand forecasting is necessary to be sure that there is adequate power supply to meet consumer's needs. While Texas has a forecasting model created by the Electricity Reliability Council of Texas (ERCOT), constant efforts are required to ensure that the model stays at the state-of-the-art and is producing the most reliable forecasts possible. This research seeks to provide improved short- and medium-term forecasting models, bringing in state-of-the-art deep learning models to compare to ERCOT’s forecasts. A model that is more accurate than ERCOT’s own …
Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia
Identification And Characterization Of Forest Fire Risk Zones Leveraging Machine Learning Methods, Joshua Balson, Matt Chinchilla, Cam Lu, Jeff Washburn, Nibhrat Lohia
SMU Data Science Review
Across the United States, record numbers of wildfires are observed costing billions of dollars in property damage, polluting the environment, and putting lives at risk. The ability of emergency management professionals, city planners, and private entities such as insurance companies to determine if an area is at higher risk of a fire breaking out has never been greater. This paper proposes a novel methodology for identifying and characterizing zones with increased risks of forest fires. Methods involving machine learning techniques use the widely available and recorded data, thus making it possible to implement the tool quickly.
Comparing Machine Learning Techniques With State-Of-The-Art Parametric Prediction Models For Predicting Soybean Traits, Susweta Ray
Department of Statistics: Dissertations, Theses, and Student Research
Soybean is a significant source of protein and oil, and also widely used as animal feed. Thus, developing lines that are superior in terms of yield, protein and oil content is important to feed the ever-growing population. As opposed to the high-cost phenotyping, genotyping is both cost and time efficient for breeders while evaluating new lines in different environments (location-year combinations) can be costly. Several Genomic prediction (GP) methods have been developed to use the marker and environment data effectively to predict the yield or other relevant phenotypic traits of crops. Our study compares a conventional GP method (GBLUP), a …