Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (188)
- Life Sciences (182)
- Computer Sciences (170)
- Bioinformatics (154)
- Biomedical Informatics (153)
-
- Statistics and Probability (85)
- Engineering (84)
- Social and Behavioral Sciences (84)
- Artificial Intelligence and Robotics (81)
- Medical Sciences (56)
- Computer Engineering (41)
- Medical Specialties (39)
- Electrical and Computer Engineering (31)
- Applied Mathematics (28)
- Databases and Information Systems (28)
- Statistical Models (28)
- Applied Statistics (27)
- Business (27)
- Public Health (27)
- Data Storage Systems (25)
- Mathematics (24)
- Theory and Algorithms (23)
- Diseases (22)
- Environmental Sciences (21)
- Systems and Communications (21)
- Oncology (18)
- Other Computer Sciences (18)
- Sociology (18)
- Institution
-
- The Texas Medical Center Library (151)
- Southern Methodist University (26)
- Old Dominion University (24)
- Universitas Negeri Malang (18)
- City University of New York (CUNY) (15)
-
- University of Central Florida (14)
- Chapman University (12)
- Embry-Riddle Aeronautical University (12)
- Purdue University (10)
- Claremont Colleges (9)
- Illinois State University (9)
- Kennesaw State University (9)
- San Jose State University (8)
- Smith College (8)
- California State University, San Bernardino (7)
- Dartmouth College (7)
- East Tennessee State University (7)
- New Jersey Institute of Technology (7)
- Southern Adventist University (7)
- Mississippi State University (6)
- Air Force Institute of Technology (5)
- California Polytechnic State University, San Luis Obispo (5)
- Clemson University (5)
- Louisiana State University (5)
- University of Kentucky (5)
- University of Louisville (5)
- University of Nevada, Las Vegas (5)
- Utah State University (5)
- West Virginia University (5)
- Bryant University (4)
- Keyword
-
- Humans (85)
- Machine learning (44)
- Machine Learning (39)
- Deep learning (21)
- Deep Learning (18)
-
- Natural Language Processing (15)
- Animals (14)
- Data Science (14)
- Data science (14)
- Electronic Health Records (14)
- Artificial Intelligence (13)
- COVID-19 (12)
- Mice (11)
- Classification (10)
- Gene Expression Profiling (10)
- Natural language processing (10)
- Algorithms (9)
- Artificial intelligence (8)
- Privacy (8)
- Analysis (7)
- Brain (7)
- Data (7)
- Genome-Wide Association Study (7)
- Genomics (7)
- NLP (7)
- Neural Networks (7)
- Neural networks (7)
- Single-Cell Analysis (7)
- Software (7)
- United States (7)
- Publication
-
- Faculty, Staff and Student Publications (146)
- SMU Data Science Review (24)
- Knowledge Engineering and Data Science (18)
- Data Science and Data Mining (14)
- Dissertations (13)
-
- Theses and Dissertations (13)
- Annual Symposium on Biomathematics and Ecology Education and Research (9)
- Electronic Theses and Dissertations (9)
- Computational and Data Sciences (PhD) Dissertations (7)
- Electronic Theses, Projects, and Dissertations (7)
- Master's Projects (7)
- Modeling, Simulation and Visualization Student Capstone Conference (7)
- Dissertations, Theses, and Capstone Projects (6)
- I-GUIDE Forum (6)
- Statistical and Data Sciences: Faculty Publications (6)
- All Dissertations (5)
- CMC Senior Theses (5)
- Campus Research Month (5)
- Dissertations and Theses (Open Access) (5)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (5)
- International Conference on Gambling & Risk Taking (5)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (4)
- College of Engineering Summer Undergraduate Research Program (4)
- College of Graduate Studies: Theses & Dissertations (4)
- Dissertations and Theses (4)
- Faculty Publications (4)
- Math Department Colloquium Series (4)
- Publications and Research (4)
- Symposium of Student Scholars (4)
- All Graduate Theses, Dissertations, and Other Capstone Projects (3)
- Publication Type
- File Type
Articles 451 - 480 of 545
Full-Text Articles in Data Science
Predicting Heart Disease Using Tree-Based Model, Emil Agbemade
Predicting Heart Disease Using Tree-Based Model, Emil Agbemade
Data Science and Data Mining
The paper presents a study on the use of machine learning algorithms for the prediction of heart disease, which is the leading cause of death worldwide. The study focuses on the use of decision tree algorithms, which have the advantage of considering a large number of risk factors. The heart disease data set was obtained from the UCI Machine Learning Repository and was analyzed using a decision tree classifier. The data set had 6 missing data points, which were deleted, leaving 279 instances for analysis. One-hot-encoding was performed on categorical variables with more than two responses. The decision tree classifier …
Machine Learning-Based Approaches For Predicting The Critical Temperature Of Superconductor, Pradip Dhakal
Machine Learning-Based Approaches For Predicting The Critical Temperature Of Superconductor, Pradip Dhakal
Data Science and Data Mining
This paper focuses on utilizing multiple linear regression, lasso regression, and extreme gradient boosting algorithms to predict the critical temperature of the superconductor. The model will be evaluated using the mean square error and adjusted R-squared values, and the best model will be recommended for future work related to this study.
Variable Selection Using Lasso And Elastic Net Regression On High Dimensional Genetic Architecture Data Of Maize Flowering Time, Pradip Dhakal
Variable Selection Using Lasso And Elastic Net Regression On High Dimensional Genetic Architecture Data Of Maize Flowering Time, Pradip Dhakal
Data Science and Data Mining
Variable selection is one of the key components in the machine learning area. This method reduces the unwanted and redundant predictors in the model, which prevents the overfitting situation. Since the model contains few significant predictors, the model is less likely to learn the trend from the noise. Further, the time to train the model reduces when we have only a few valuable variables.
Silent Agony: Automated Detection Of Ethnic And Religious Cyberbullying Using Machine Learning, Emil Agbemade
Silent Agony: Automated Detection Of Ethnic And Religious Cyberbullying Using Machine Learning, Emil Agbemade
Data Science and Data Mining
The use of electronic mobile devices, social media, and networking websites has increased tremendously in recent years. Despite the advantages of these systems, such as exchanging ideas and information, being sociable, and providing entertainment, users may encounter adverse behaviors like toxicity, bullying, extremism, and cruelty. The prevalence of such behaviors has grown significantly in cyberspace, posing a threat to individuals and communities. To address this issue, there is a high demand for automated cyberbullying detection systems. Machine learning algorithms have been widely used to build such systems by classifying and detecting cyberbullying. In this study, we employed popular machine learning …
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Theses and Dissertations--Computer Science
As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally - using only a measure of task performance as feedback--can violate societal norms for acceptable behavior or cause harm. Consequently, it becomes necessary to prioritize task performance and ensure that AI actions do not have detrimental effects. Value alignment is a property of intelligent agents, wherein they solely pursue goals and activities that are non-harmful and beneficial to humans. Current approaches to value alignment largely depend on imitation learning or learning from demonstration methods. However, the dynamic nature …
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Theses and Dissertations--Statistics
Neural networks have experienced widespread adoption and have become integral in cutting-edge domains like computer vision, natural language processing, and various contemporary fields. However, addressing the statistical aspects of neural networks has been a persistent challenge, with limited satisfactory results. In my research, I focused on exploring statistical intervals applied to neural networks, specifically confidence intervals and tolerance intervals. I employed variance estimation methods, such as direct estimation and resampling, to assess neural networks and their performance under outlier scenarios. Remarkably, when outliers were present, the resampling method with infinitesimal jackknife estimation yielded confidence intervals that closely aligned with nominal …
High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang
High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang
Theses and Dissertations--Statistics
This dissertation focuses on the problem of high dimensional data analysis, which arises in many fields including genomics, finance, and social sciences. In such settings, the number of features or variables is much larger than the number of observations, posing significant challenges to traditional statistical methods.
To address these challenges, this dissertation proposes novel methods for variable screening and inference. The first part of the dissertation focuses on variable screening, which aims to identify a subset of important variables that are strongly associated with the response variable. Specifically, we propose a robust nonparametric screening method to effectively select the predictors …
Clustering Hospital Performance Using Group-Based Multi-Trajectory Modeling With Singular Bayesian Information Criterion, Gaixin Du
Theses and Dissertations--Epidemiology and Biostatistics
Hospital performance is complex and patient-experience oriented. Currently, the Centers for Medicare and Medicaid Services (CMS) evaluate hospitals yearly with a single score of one to five ("Star Rating") using composite measures from five domains. However, a single composite score cannot fully describe it, and alternative measures should be considered. Healthcare quality improvement needs long-term data to validate effectiveness. Group-based multi-trajectory modeling (GBMTM) estimates probabilities of latent group membership based on longitudinal profiles from multiple outcomes. We use GBMTM to identify groups of hospitals with similar performance in SAS PROC TRAJ.
We downloaded Medicare-eligible hospitals (N=5,111) that provided patient care …
The Gdpr And Uk Gdpr And Its Impact On Us Academic Institutions, Leila Halawi, Alpesh Makwana
The Gdpr And Uk Gdpr And Its Impact On Us Academic Institutions, Leila Halawi, Alpesh Makwana
Publications
This research paper delves into the implications of the General Data Protection Regulation (GDPR) and the United Kingdom (UK) GDPR on academic institutions, shedding light on their significance for organizations and educational establishments handling data from individuals in the European Union (EU) and the UK. Non-compliance with these regulations can lead to substantial penalties. The study focuses specifically on US Higher Education and presents actionable measures that institutions can adopt to enhance compliance, fortify data protection, and safeguard the privacy of individuals.
Covidanno, Covid-19 Annotation In Human, Yuzhou Feng, Mengyuan Yang, Zhiwei Fan, Weiling Zhao, Pora Kim, Xiaobo Zhou
Covidanno, Covid-19 Annotation In Human, Yuzhou Feng, Mengyuan Yang, Zhiwei Fan, Weiling Zhao, Pora Kim, Xiaobo Zhou
Faculty, Staff and Student Publications
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the etiologic agent of coronavirus disease 19 (COVID-19), has caused a global health crisis. Despite ongoing efforts to treat patients, there is no universal prevention or cure available. One of the feasible approaches will be identifying the key genes from SARS-CoV-2-infected cells. SARS-CoV-2-infected in vitro model, allows easy control of the experimental conditions, obtaining reproducible results, and monitoring of infection progression. Currently, accumulating RNA-seq data from SARS-CoV-2 in vitro models urgently needs systematic translation and interpretation. To fill this gap, we built COVIDanno, COVID-19 annotation in humans, available at http://biomedbdc.wchscu.cn/COVIDanno/. The aim …
Syllable-Pbwt For Space-Efficient Haplotype Long-Match Query, Victor Wang, Ardalan Naseri, Shaojie Zhang, Degui Zhi
Syllable-Pbwt For Space-Efficient Haplotype Long-Match Query, Victor Wang, Ardalan Naseri, Shaojie Zhang, Degui Zhi
Faculty, Staff and Student Publications
MOTIVATION: The positional Burrows-Wheeler transform (PBWT) has led to tremendous strides in haplotype matching on biobank-scale data. For genetic genealogical search, PBWT-based methods have optimized the asymptotic runtime of finding long matches between a query haplotype and a predefined panel of haplotypes. However, to enable fast query searches, the full-sized panel and PBWT data structures must be kept in memory, preventing existing algorithms from scaling up to modern biobank panels consisting of millions of haplotypes. In this work, we propose a space-efficient variation of PBWT named Syllable-PBWT, which divides every haplotype into syllables, builds the PBWT positional prefix arrays on …
Use Gpt-J Prompt Generation With Roberta For Ner Models On Diagnosis Extraction Of Periodontal Diagnosis From Electronic Dental Records, Yao-Shun Chuang, Xiaoqian Jiang, Chun-Teh Lee, Ryan Brandon, Duong Tran, Oluwabunmi Tokede, Muhammad F Walji
Use Gpt-J Prompt Generation With Roberta For Ner Models On Diagnosis Extraction Of Periodontal Diagnosis From Electronic Dental Records, Yao-Shun Chuang, Xiaoqian Jiang, Chun-Teh Lee, Ryan Brandon, Duong Tran, Oluwabunmi Tokede, Muhammad F Walji
Faculty, Staff and Student Publications
This study explored the usability of prompt generation on named entity recognition (NER) tasks and the performance in different settings of the prompt. The prompt generation by GPT-J models was utilized to directly test the gold standard as well as to generate the seed and further fed to the RoBERTa model with the spaCy package. In the direct test, a lower ratio of negative examples with higher numbers of examples in prompt achieved the best results with a F1 score of 0.72. The performance revealed consistency, 0.92-0.97 in the F1 score, in all settings after training with the RoBERTa model. …
Towards Fair Patient-Trial Matching Via Patient-Criterion Level Fairness Constraint, Chia-Yuan Chang, Jiayi Yuan, Sirui Ding, Qiaoyu Tan, Kai Zhang, Xiaoqian Jiang, Xia Hu, Na Zou
Towards Fair Patient-Trial Matching Via Patient-Criterion Level Fairness Constraint, Chia-Yuan Chang, Jiayi Yuan, Sirui Ding, Qiaoyu Tan, Kai Zhang, Xiaoqian Jiang, Xia Hu, Na Zou
Faculty, Staff and Student Publications
Clinical trials are indispensable in developing new treatments, but they face obstacles in patient recruitment and retention, hindering the enrollment of necessary participants. To tackle these challenges, deep learning frameworks have been created to match patients to trials. These frameworks calculate the similarity between patients and clinical trial eligibility criteria, considering the discrepancy between inclusion and exclusion criteria. Recent studies have shown that these frameworks outperform earlier approaches. However, deep learning models may raise fairness issues in patient-trial matching when certain sensitive groups of individuals are underrepresented in clinical trials, leading to incomplete or inaccurate data and potential harm. To …
Maximizing Productivity And Quality In Senior Thesis Writing With Artificial Intelligence And Natural Language Processing Driven Tools, Lauren Leadbetter
Maximizing Productivity And Quality In Senior Thesis Writing With Artificial Intelligence And Natural Language Processing Driven Tools, Lauren Leadbetter
CMC Senior Theses
This project is a Python program designed to generate a senior thesis on a user-
inputted topic using natural language processing techniques. The program takes in a
topic from the user and then uses OpenAI API to deploy text models for text genera-
tion and evaluation, such as GPT-3 and Davinci-003. The resulting output is in .tex
format and includes a first-draft outline and paper, followed by self-generated assessment, with scoring, revisions, and feedback comments instructing manual revisions.
This submission is a sample using one available model of the project, meant to
demonstrate it’s functionality and limitations. Further model versions …
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
Theses and Dissertations--Computer Science
We introduce a novel approach for learning behaviors using human-provided feedback that is subject to systematic bias. Our method, known as BASIL, models the feedback signal as a combination of a heuristic evaluation of an action's utility and a probabilistically-drawn bias value, characterized by unknown parameters. We present both the general framework for our technique and specific algorithms for biases drawn from a normal distribution. We evaluate our approach across various environments and tasks, comparing it to interactive and non-interactive machine learning methods, including deep learning techniques, using human trainers and a synthetic oracle with feedback distorted to varying degrees. …
Should Academia Thrive For Research Citation In Policy? A Case Study On Five Universities In Illinois., Minhaz Suleman Ibrahim Patel
Should Academia Thrive For Research Citation In Policy? A Case Study On Five Universities In Illinois., Minhaz Suleman Ibrahim Patel
CURE Proceedings
Academics and policymakers are seen as operating separately, which limits the potential impact of research on society. The influence of university research on policy documents is frequently underestimated, given that cutting-edge research is being conducted at universities. Therefore, it is crucial to unveil the role of academic research in fostering evidence-driven policymaking across various public service domains. In this study, we conducted an in-depth exploratory data analysis and statistical summarization to comprehensively understand the level of academic research present in policy documents. We chose five public universities from the state of Illinois and collected research and policy citation data for …
The Shortfalls Of Vulnerability Indexes For Public Health Decision-Making In The Face Of Emergent Crises: The Case Of Covid-19 Vaccine Uptake In Virginia, Lydia Cleveland Sa, Erika Frydenlund
The Shortfalls Of Vulnerability Indexes For Public Health Decision-Making In The Face Of Emergent Crises: The Case Of Covid-19 Vaccine Uptake In Virginia, Lydia Cleveland Sa, Erika Frydenlund
VMASC Publications
Equitable and effective vaccine uptake is a key issue in addressing COVID-19. To achieve this, we must comprehensively characterize the context-specific socio-behavioral and structural determinants of vaccine uptake. However, to quickly focus public health interventions, state agencies and planners often rely on already existing indexes of "vulnerability." Many such "vulnerability indexes" exist and become benchmarks for targeting interventions in wide ranging scenarios, but they vary considerably in the factors and themes that they cover. Some are even uncritical of the use of the word "vulnerable," which should take on different meanings in different contexts. The objective of this study is …
Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede
Evaluation Of Edison's Data Science Competency Framework Through A Comparative Literature Analysis, Karl R. B. Schmitt, Linda Clark, Katherine M. Kinnaird, Ruth E. H. Wertz, Björn Sandstede
Statistical and Data Sciences: Faculty Publications
During the emergence of Data Science as a distinct discipline, discussions of what exactly constitutes Data Science have been a source of contention, with no clear resolution. These disagreements have been exacerbated by the lack of a clear single disciplinary 'parent.' Many early efforts at defining curricula and courses exist, with the EDISON Project's Data Science Framework (EDISON-DSF) from the European Union being the most complete. The EDISON-DSF includes both a Data Science Body of Knowledge (DS-BoK) and Competency Framework (CF-DS). This paper takes a critical look at how EDISON's CF-DS compares to recent work and other published curricular or …
Deep Learning-Based Technique For The Perception Of The Cervical Cancer, Aya Haraz, Hossam El-Din Moustafa, Abeer Twakol Khaleel, Ahmed H. Eltanboly
Deep Learning-Based Technique For The Perception Of The Cervical Cancer, Aya Haraz, Hossam El-Din Moustafa, Abeer Twakol Khaleel, Ahmed H. Eltanboly
Mansoura Engineering Journal
In third-world countries, cervical cancer is the most prevalent and leading cause of death. It is affected by a variety of factors, including smoking, poor nutritional status, immunological inadequacy, and prolonged use of contraception. The Pap smear test, which is intended to prevent cervical cancer, finds preneoplastic changes in cervical epithelial cells. This study framework classified cervical cancer cells from Pap smears into five specified cell types using machine learning-based classification algorithms. The SIPaKMeD database is used in this investigation. This public dataset, which was manually cropped from 966 cluster cell images taken from Pap smear slides, has 4045 isolated …
Blockchain And Puf-Based Secure Key Establishment Protocol For Cross-Domain Digital Twins In Industrial Internet Of Things Architecture, Khalid Mahmood, Salman Shamshad, Muhammad Asad Saleem, Rupak Kharel, Ashok Kumar Das, Sachin Shetty, Joel J. P. C. Rodrigues
Blockchain And Puf-Based Secure Key Establishment Protocol For Cross-Domain Digital Twins In Industrial Internet Of Things Architecture, Khalid Mahmood, Salman Shamshad, Muhammad Asad Saleem, Rupak Kharel, Ashok Kumar Das, Sachin Shetty, Joel J. P. C. Rodrigues
VMASC Publications
Introduction:: The Industrial Internet of Things (IIoT) is a technology that connects devices to collect data and conduct in-depth analysis to provide value-added services to industries. The integration of the physical and digital domains is crucial for unlocking the full potential of the IIoT, and digital twins can facilitate this integration by providing a virtual representation of real-world entities.
Objectives:: By combining digital twins with the IIoT, industries can simulate, predict, and control physical behaviors, enabling them to achieve broader value and support industry 4.0 and 5.0. Constituents of cooperative IIoT domains tend to interact and collaborate during their complicated …
Lessons Learned From Interdisciplinary Efforts To Combat Covid-19 Misinformation: Development Of Agile Integrative Methods From Behavioral Science, Data Science, And Implementation Science, Sahiti Myneni, Paula Cuccaro, Sarah Montgomery, Vivek Pakanati, Jinni Tang, Tavleen Singh, Olivia Dominguez, Trevor Cohen, Belinda Reininger, Lara S Savas, Maria E Fernandez
Lessons Learned From Interdisciplinary Efforts To Combat Covid-19 Misinformation: Development Of Agile Integrative Methods From Behavioral Science, Data Science, And Implementation Science, Sahiti Myneni, Paula Cuccaro, Sarah Montgomery, Vivek Pakanati, Jinni Tang, Tavleen Singh, Olivia Dominguez, Trevor Cohen, Belinda Reininger, Lara S Savas, Maria E Fernandez
Faculty, Staff and Student Publications
BACKGROUND: Despite increasing awareness about and advances in addressing social media misinformation, the free flow of false COVID-19 information has continued, affecting individuals' preventive behaviors, including masking, testing, and vaccine uptake.
OBJECTIVE: In this paper, we describe our multidisciplinary efforts with a specific focus on methods to (1) gather community needs, (2) develop interventions, and (3) conduct large-scale agile and rapid community assessments to examine and combat COVID-19 misinformation.
METHODS: We used the Intervention Mapping framework to perform community needs assessment and develop theory-informed interventions. To supplement these rapid and responsive efforts through large-scale online social listening, we developed a …
A Query Engine For Self-Controlled Case Series, With An Application To Covid-19 Ehr Data, Xiaojin Li, Yan Huang, Licong Cui, Guo-Qiang Zhang
A Query Engine For Self-Controlled Case Series, With An Application To Covid-19 Ehr Data, Xiaojin Li, Yan Huang, Licong Cui, Guo-Qiang Zhang
Faculty, Staff and Student Publications
Self-controlled case series (SCCS) is a statistical method in epidemiological study design that uses individuals as their own controls, with comparisons made within the same individuals at different time points of observation. SCCS has been applied in settings where it is difficult to identify comparison or control groups. To provide computational support for SCCS, we introduce a query engine called Self-Controlled Case Query (SCCQ) and use it to extract cohorts of self-controlled case series from a large-scale COVID-19 Electronic Health Records (EHR) dataset. Visual summary of the queried population through the R-Shiny visualization framework offers SCCQ's query result dashboard to …
Federated Learning Based Futuristic Biomedical Big-Data Analysis And Standardization, Afifa Salsabil Fathima, Syed Muzamil Basha, Syed Thouheed Ahmed, Sandeep Kumar Mathivanan, Sukumar Rajendran, Saurav Mallik, Zhongming Zhao
Federated Learning Based Futuristic Biomedical Big-Data Analysis And Standardization, Afifa Salsabil Fathima, Syed Muzamil Basha, Syed Thouheed Ahmed, Sandeep Kumar Mathivanan, Sukumar Rajendran, Saurav Mallik, Zhongming Zhao
Faculty, Staff and Student Publications
Medical data processing and analytics exert significant influence in furnishing dependable decision support for prospective biomedical applications. Given the sensitive nature of medical data, specialized techniques and frameworks tailored for application-centric processing are imperative. This article presents a conceptualization for the analysis and uniformitarian of datasets through the implementation of Federated Learning (FL). The realm of medical big data stems from diverse origins, necessitating the delineation of data provenance and attribute paradigms to facilitate feature extraction and dependency assessment. The architecture governing the data collection framework is intricately linked to remote data transmission, thereby engendering efficient customization oversight. The operational …
Warehouses In The Inland Empire: Displacing Land And Life, Katherine Gelsey
Warehouses In The Inland Empire: Displacing Land And Life, Katherine Gelsey
Pomona Senior Theses
The Inland Empire in Southern California embodies unique spatial and social configurations as a consequence of how settler colonialism has manifested locally in the region since the Spanish Mission Period. This work uses GIS software to estimate patterns of land conversion for residential, agricultural, and warehouse land from 2012 to 2022. Preliminary analysis suggests that thousands of people have been displaced by warehouse expansion over the ten-year period. In the twenty-first century, the Southern California logistics industry continues processes of land dispossession and racialized labor exploitation through displacing agricultural and residential land, exposing disproportionately low-income Black and Latine communities living …
Applications Of Transfer Learning From Malicious To Vulnerable Binaries, Sean Patrick Mcnulty
Applications Of Transfer Learning From Malicious To Vulnerable Binaries, Sean Patrick Mcnulty
Graduate Student Theses, Dissertations, & Professional Papers
Malware detection and vulnerability detection are important cybersecurity tasks. Previous research has successfully applied a variety of machine learning methods to both. However, despite their potential synergies, previous research has yet to unite these two tasks. Given the recent success of transfer learning in many domains, such as language modeling and image recognition, this thesis investigated the use of transfer learning to improve vulnerability detection. Specifically, we pre-trained a series of models to detect malicious binaries and used the weights from those models to kickstart the detection of vulnerable binaries. In our study, we also investigated five different data representations …
Development Of A Data Science Curriculum For An Engineering Technology Program, Salih Sarp, Murat Kuzlu, Otilia Popescu, Vukica M. Jovanovic, Zafer Acar
Development Of A Data Science Curriculum For An Engineering Technology Program, Salih Sarp, Murat Kuzlu, Otilia Popescu, Vukica M. Jovanovic, Zafer Acar
Engineering Technology Faculty Publications
Data science has gained the attention of various industries, educators, parents, and students thinking about their future careers. Statistics departments have traditionally offered data science courses for a long time. The main objective of these courses is to examine the fundamental concepts and theories. However, teaching data science courses has also expanded to other disciplines due to the vast amount of data being collected by numerous modern applications. Also, someone needs to learn how to collect and process data, especially from industrial devices, because of the recent development of Internet of Things (IoT) technologies. Hence, integrating data science into the …
Automating Intersection Marking Data Collection And Condition Assessment At Scale With An Artificial Intelligence-Powered System, Kun Xie, Huiming Sun, Xiaomeng Dong, Hong Yang, Hongkai Yu
Automating Intersection Marking Data Collection And Condition Assessment At Scale With An Artificial Intelligence-Powered System, Kun Xie, Huiming Sun, Xiaomeng Dong, Hong Yang, Hongkai Yu
Civil & Environmental Engineering Faculty Publications
Intersection markings play a vital role in providing road users with guidance and information. The conditions of intersection markings will be gradually degrading due to vehicular traffic, rain, and/or snowplowing. Degraded markings can confuse drivers, leading to increased risk of traffic crashes. Timely obtaining high-quality information of intersection markings lays a foundation for making informed decisions in safety management and maintenance prioritization. However, current labor-intensive and high-cost data collection practices make it very challenging to gather intersection data on a large scale. This paper develops an automated system to intelligently detect intersection markings and to assess their degradation conditions with …
Biodiversity Of Philippine Marine Fishes: A Dna Barcode Reference Library Based On Voucher Specimens, Katherine E. Bemis, Matthew G. Girard, Mudjekeewis D. Santos, Kent E. Carpenter, Jonathan R. Deeds, Diane E. Pitassy, Nicko Amor L. Flores, Elizabeth S. Hunter, Amy C. Driskell, Kenneth S. Macdonald Iii, Lee A. Weigt, Jeffrey T. Williams
Biodiversity Of Philippine Marine Fishes: A Dna Barcode Reference Library Based On Voucher Specimens, Katherine E. Bemis, Matthew G. Girard, Mudjekeewis D. Santos, Kent E. Carpenter, Jonathan R. Deeds, Diane E. Pitassy, Nicko Amor L. Flores, Elizabeth S. Hunter, Amy C. Driskell, Kenneth S. Macdonald Iii, Lee A. Weigt, Jeffrey T. Williams
Biological Sciences Faculty Publications
Accurate identification of fishes is essential for understanding their biology and to ensure food safety for consumers. DNA barcoding is an important tool because it can verify identifications of both whole and processed fishes that have had key morphological characters removed (e.g., filets, fish meal); however, DNA reference libraries are incomplete, and public repositories for sequence data contain incorrectly identified sequences. During a nine-year sampling program in the Philippines, a global biodiversity hotspot for marine fishes, we developed a verified reference library of cytochrome c oxidase subunit I (COI) sequences for 2,525 specimens representing 984 species. Specimens were primarily purchased …
A Data Analysis On Mass Shootings In Amercia, James Hinkle, Patrick Mccool
A Data Analysis On Mass Shootings In Amercia, James Hinkle, Patrick Mccool
Capstone Showcase
Mass shootings in America have been a recurring issue for years. In this project, we examine mass shootings that have occurred in the United States from 1966 to 2022. Through exploratory data analyses, we explore patterns and trends in shooting events, as well as various patterns in shooters, such as their mental health status, relationship status, social media usage, evidence of trauma in adulthood, and ongoing stressors during the time of the shooting. We also utilize natural language processing (NLP) tools to analyze text information in the dataset, such as the shooters' school performance, community involvement, and past signs of …
Quantifying The Carbon Stored And Sequestered By The Trees On Pomona College’S Campus, Paola A. Giron-Carson
Quantifying The Carbon Stored And Sequestered By The Trees On Pomona College’S Campus, Paola A. Giron-Carson
Scripps Senior Theses
We are experiencing a climate crisis that must be confronted with strategic mitigation. Pomona College contributes to the climate crisis through its emissions for which there is a baseline record. However there is no baseline record of the climate mitigation currently performed by the trees on Pomona’s campus through carbon storage. This study seeks to determine a current baseline quantity of carbon stored and sequestrated by Pomona’s trees as well as possible courses of climate mitigation for Pomona College to take. Initial information gathering was conducted through interviews with several stakeholders. This study was conducted using data collected prior to …