Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (1157)
- Medicine and Health Sciences (780)
- Life Sciences (765)
- Bioinformatics (568)
- Statistics and Probability (550)
-
- Biomedical Informatics (530)
- Engineering (528)
- Artificial Intelligence and Robotics (526)
- Social and Behavioral Sciences (520)
- Databases and Information Systems (212)
- Computer Engineering (208)
- Electrical and Computer Engineering (204)
- Applied Statistics (194)
- Medical Sciences (190)
- Business (189)
- Statistical Models (181)
- Applied Mathematics (175)
- Medical Specialties (173)
- Environmental Sciences (149)
- Theory and Algorithms (149)
- Mathematics (144)
- Other Computer Sciences (127)
- Data Storage Systems (123)
- Systems and Communications (120)
- Numerical Analysis and Scientific Computing (116)
- Public Health (116)
- Public Affairs, Public Policy and Public Administration (109)
- Statistical Methodology (109)
- Institution
-
- The Texas Medical Center Library (523)
- Old Dominion University (173)
- Southern Methodist University (144)
- Universitas Negeri Malang (113)
- City University of New York (CUNY) (101)
-
- CCT College Dublin (91)
- Chapman University (67)
- Kennesaw State University (63)
- University of Central Florida (62)
- Smith College (60)
- Air Force Institute of Technology (57)
- Embry-Riddle Aeronautical University (52)
- Singapore Management University (45)
- University of Arkansas, Fayetteville (45)
- Chinese Academy of Sciences (44)
- Purdue University (44)
- California Polytechnic State University, San Luis Obispo (39)
- Technological University Dublin (39)
- Illinois State University (38)
- University of Kentucky (38)
- University of Nebraska - Lincoln (38)
- New Jersey Institute of Technology (37)
- West Virginia University (37)
- Claremont Colleges (36)
- Virginia Commonwealth University (35)
- Clemson University (32)
- Dartmouth College (31)
- University of Texas at Arlington (27)
- East Tennessee State University (26)
- Minnesota State University, Mankato (26)
- Keyword
-
- Humans (278)
- Machine learning (241)
- Machine Learning (218)
- Deep learning (115)
- Computer Science (107)
-
- Deep Learning (94)
- Artificial Intelligence (65)
- Data science (58)
- Data Science (57)
- Natural Language Processing (56)
- COVID-19 (55)
- Artificial intelligence (53)
- Female (52)
- Male (50)
- Classification (49)
- Natural language processing (47)
- Animals (41)
- Data (41)
- Electronic Health Records (41)
- Neural Networks (40)
- Algorithms (38)
- Big data (37)
- Data mining (37)
- Statistics (36)
- Clustering (32)
- Computer science (31)
- Adult (30)
- NLP (30)
- Neural networks (30)
- Random Forest (30)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (508)
- SMU Data Science Review (124)
- Knowledge Engineering and Data Science (113)
- Theses and Dissertations (111)
- ICT (91)
-
- Data Science and Data Mining (53)
- Dissertations (53)
- Statistical and Data Sciences: Faculty Publications (53)
- Electronic Theses and Dissertations (49)
- Dissertations, Theses, and Capstone Projects (45)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (44)
- Research Collection School Of Computing and Information Systems (37)
- Master's Theses (35)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (34)
- Data Science Undergraduate Honors Theses (31)
- Annual Symposium on Biomathematics and Ecology Education and Research (30)
- Computer Science Faculty Publications (30)
- Publications and Research (30)
- Computational and Data Sciences (PhD) Dissertations (25)
- All Graduate Theses, Dissertations, and Other Capstone Projects (24)
- Symposium of Student Scholars (24)
- All Dissertations (23)
- Articles (23)
- Electrical & Computer Engineering Faculty Publications (22)
- CBN Journal of Applied Statistics (JAS) (21)
- College of Graduate Studies: Theses & Dissertations (20)
- CMC Senior Theses (19)
- Electronic Theses, Projects, and Dissertations (19)
- Theses (19)
- Faculty Publications (18)
- Publication Type
- File Type
Articles 2941 - 2970 of 3244
Full-Text Articles in Data Science
The Diffusion Of Ict For Corruption Detectionin Open Government Data, Darusalam Darusalam, Jamaliah Said, Normah Omar, Marijn Janssen, Kazi Sohag
The Diffusion Of Ict For Corruption Detectionin Open Government Data, Darusalam Darusalam, Jamaliah Said, Normah Omar, Marijn Janssen, Kazi Sohag
Knowledge Engineering and Data Science
Corruption occurs in many places within the government. To tackle the issue, open data can be used as one of the tools in creating more insight into the government. The premise of this paper is to support the notion that data opening can bring up new ways of fighting corruption. The current paper aimed at investigating how open data can be employed to detect corruption. This open data is trivial due to challenges like information asymmetry among stakeholders, data might only be opened partly, different sources of data need to be combined, and data might not be easy to use, …
Adam Optimization Algorithmfor Wide And Deep Neural Network, Imran Khan Mohd Jais, Amelia Ritahani Ismail
Adam Optimization Algorithmfor Wide And Deep Neural Network, Imran Khan Mohd Jais, Amelia Ritahani Ismail
Knowledge Engineering and Data Science
The objective of this research is to evaluate the effects of Adam when used together with a wide and deep neural network. The dataset used was a diagnostic breast cancer dataset taken from UCI Machine Learning. Then, the dataset was fed into a conventional neural network for a benchmark test. Afterwards, the dataset was fed into the wide and deep neural network with and without Adam. It was found that there were improvements in the result of the wide and deep network with Adam. In conclusion, Adam is able to improve the performance of a wide and deep neural network.
Selection Of Marine Security Policyusing Fuzzy-Ahp Topsis Hybrid Approach, Hozairi Hozairi, Buhari Buhari, Heru Lumaksono, Marcus Tukan
Selection Of Marine Security Policyusing Fuzzy-Ahp Topsis Hybrid Approach, Hozairi Hozairi, Buhari Buhari, Heru Lumaksono, Marcus Tukan
Knowledge Engineering and Data Science
The research was focused on the integration of Fuzzy set theory with Analytic Hierarchy Process (AHP) and Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) to choose the optimum maritime security policy to achieve Indonesia recognition as the world's maritime axis. The method used is AHP with fuzzy based enhancement. Here, the weight of each criterion is calculated to overcome the criticism of the scale of unbalanced rating, uncertainty, and inaccuracy in the pairwise of comparison process. The best recommendation for Indonesian maritime policies is multi task single agency which is greatly infuenced by several factors such as …
High Dimensional Data Clustering Using Self-Organized Map, Ruth Ema Febrita, Wayan Firdaus Mahmudy, Aji Prasetya Wibawa
High Dimensional Data Clustering Using Self-Organized Map, Ruth Ema Febrita, Wayan Firdaus Mahmudy, Aji Prasetya Wibawa
Knowledge Engineering and Data Science
As the population grows and e economic development, houses could be one of basic needs of every family. Therefore, housing investment has promising value in the future. This research implements the Self-Organized Map (SOM) algorithm to cluster house data for providing several house groups based on the various features. K-means is used as the baseline of the proposed approach. SOM has higher silhouette coefficient (0.4367) compared to its comparison (0.236). Thus, this method outperforms k-means in terms of visualizing high-dimensional data cluster. It is also better in the cluster formation and regulating the data distribution.
Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan
Statistical Machine Learning Methods For Mining Spatial And Temporal Data, Fei Tan
Dissertations
Spatial and temporal dependencies are ubiquitous properties of data in numerous domains. The popularity of spatial and temporal data mining has thus grown with the increasing prevalence of massive data. The presence of spatial and temporal attributes not only provides complementary useful perspectives, but also poses new challenges to the representation and integration into the learning procedure. In this dissertation, the involved spatial and temporal dependencies are explored with three genres: sample-wise, feature-wise, and target-wise. A family of novel methodologies is developed accordingly for the dependency representation in respective scenarios.
First, dependencies among discrete, continuous and repeated observations are studied …
On Properties Of Distance-Based Entropies On Fullerene Graphs, Modjtaba Ghorbani, Matthias Dehmer, Mina Rajabi-Parsa, Abbe Mowshowitz, Frank Emmert-Streib
On Properties Of Distance-Based Entropies On Fullerene Graphs, Modjtaba Ghorbani, Matthias Dehmer, Mina Rajabi-Parsa, Abbe Mowshowitz, Frank Emmert-Streib
Publications and Research
In this paper, we study several distance-based entropy measures on fullerene graphs. These include the topological information content of a graph Ia(G), a degree-based entropy measure, the eccentric-entropy Ifs(G), the Hosoya entropy H(G) and, finally, the radial centric information entropy Hecc. We compare these measures on two infinite classes of fullerene graphs denoted by A12n+4 and B12n+6. We have chosen these measures as they are easily computable and capture meaningful graph properties. To demonstrate the utility of these measures, we investigate the Pearson correlation between them on the fullerene graphs.
Watersheds For Semi-Supervised Classification, Aditya Challa, Sravan Danda, B. S.Daya Sagar, Laurent Najman
Watersheds For Semi-Supervised Classification, Aditya Challa, Sravan Danda, B. S.Daya Sagar, Laurent Najman
Journal Articles
Watershed technique from mathematical morphology (MM) is one of the most widely used operators for image segmentation. Recently watersheds are adapted to edge weighted graphs, allowing for wider applicability. However, a few questions remain to be answered - How do the boundaries of the watershed operator behave? Which loss function does the watershed operator optimize? How does watershed operator relate with existing ideas from machine learning. In this letter, a framework is developed, which allows one to answer these questions. This is achieved by generalizing the maximum margin principle to maximum margin partition and proposing a generic solution, morphMedian, resulting …
Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett
Do Misperceptions Of Peer Drinking Influence Personal Drinking Behavior? Results From A Complete Social Network Of First-Year College Students, Melissa J. Cox, Angelo M. Dibello, Matthew K. Meisel, Miles Q. Ott, Shannon R. Kenney, Melissa A. Clark, Nancy P. Barnett
Statistical and Data Sciences: Faculty Publications
This study considered the influence of misperceptions of typical versus self-identified important peers' heavy drinking on personal heavy drinking intentions and frequency utilizing data from a complete social network of college students. The study sample included data from 1,313 students (44% male, 57% White, 15% Hispanic/Latinx) collected during the fall and spring semesters of their freshman year. Students provided perceived heavy drinking frequency for a typical student peer and up to 10 identified important peers. Personal past-month heavy drinking frequency was assessed for all participants at both time points. By comparing actual with perceived heavy drinking frequencies, measures of misperceptions …
Forecasting The Number Of Monthly Active Facebook And Twitter Worldwide Users Using Arma Model, Qasem Abu Al-Haija, Qian Mao, Kamal Al Nasr
Forecasting The Number Of Monthly Active Facebook And Twitter Worldwide Users Using Arma Model, Qasem Abu Al-Haija, Qian Mao, Kamal Al Nasr
Computer Science Faculty Research
In this study, an Auto-Regressive Moving Average (ARMA) Model with optimal order has been developed to estimate and forecast the short term future numbers of the monthly active Facebook and Twitter worldwide users. In order to pickup the optimal estimation order, we analyzed the model order vs. the corresponding model error in terms of final prediction error. The simulation results showed that the optimal model order to estimate the given Facebook and Twitter time series are ARMA[5, 5] and ARMA[3, 3], respectively, since they correspond to the minimum acceptable prediction error values. Besides, the optimal models recorded a high-level of …
A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer
A Grammar For Reproducible And Painless Extract-Transform-Load Operations On Medium Data, Benjamin S. Baumer
Statistical and Data Sciences: Faculty Publications
Many interesting datasets available on the Internet are of a medium size—too big to fit into a personal computer’s memory, but not so large that they would not fit comfortably on its hard disk. In the coming years, datasets of this magnitude will inform vital research in a wide array of application domains. However, due to a variety of constraints they are cumbersome to ingest, wrangle, analyze, and share in a reproducible fashion. These obstructions hamper thorough peer-review and thus disrupt the forward progress of science. We propose a predictable and pipeable framework for R (the state-of-the-art statistical computing environment) …
Exploring And Visualizing Household Electricity Consumption Patterns In Singapore: A Geospatial Analytics Approach, Yong Ying Tan, Tin Seong Kam
Exploring And Visualizing Household Electricity Consumption Patterns In Singapore: A Geospatial Analytics Approach, Yong Ying Tan, Tin Seong Kam
Research Collection School Of Computing and Information Systems
Despite being a small country-state, electricity consumption in Singa-pore is said to be non-homogeneous, as exploratory data analysis showed that the distributions of electricity consumption differ across and within administrative boundaries and dwelling types. Local indicators of spatial association (LISA) were calculated for public housing postal codes using June 2016 data to discover local clusters of households based on electricity consumption patterns. A detailed walkthrough of the analytical process is outlined to describe the R packages and framework used in the R environment. The LISA results are visualized on three levels: country level, regional level and planning subzone level. At …
Bridge Deck Delamination Segmentation Based On Aerial Thermography Through Regularized Grayscale Morphological Reconstruction And Gradient Statistics, Chongsheng Cheng, Zhexiong Shang, Zhigang Shen
Bridge Deck Delamination Segmentation Based On Aerial Thermography Through Regularized Grayscale Morphological Reconstruction And Gradient Statistics, Chongsheng Cheng, Zhexiong Shang, Zhigang Shen
Department of Construction Engineering and Management: Faculty Publications
Environmental and surface texture-induced temperature variation across the bridge deck is a major source of errors in delamination detection through thermography. This type of external noise poises a significant challenge for conventional quantitative methods such as global thresholding and k-means clustering. An iterative top-down approach is proposed for delamination segmentation based on grayscale morphological reconstruction. A weight-decay function was used to regularize the reconstruction for regional maxima extraction. The mean and coefficient of variation of temperature gradient estimated from delamination boundaries were used for discrimination. The proposed approach was tested on a lab experiment and an in-service bridge deck. The …
Thermographic Laplacian-Pyramid Filtering To Enhance Delamination Detection In Concrete Structure, Chongsheng Cheng, Ri Na, Zhigang Shen
Thermographic Laplacian-Pyramid Filtering To Enhance Delamination Detection In Concrete Structure, Chongsheng Cheng, Ri Na, Zhigang Shen
Department of Construction Engineering and Management: Faculty Publications
Despite decades of efforts using thermography to detect delamination in concrete decks, challenges still exist in removing environmental noise from thermal images. The performance of conventional temperature-contrast approaches can be significantly limited by environment-induced non-uniform temperature distribution across imaging spaces. Time-series based methodologies were found robust to spatial temperature non-uniformity but requires extended period to collect data. A new empirical image filtering method is introduced in this paper to enhance the delamination detection using blob detection method that originated from computer vison. The proposed method employs a Laplacian of Gaussian filter to achieve multi-scale detection of abnormal thermal patterns by …
Streaming Feature Grouping And Selection (Sfgs) For Big Data Classification, Noura Helal Hamad Al Nuaimi
Streaming Feature Grouping And Selection (Sfgs) For Big Data Classification, Noura Helal Hamad Al Nuaimi
Dissertations
Real-time data has always been an essential element for organizations when the quickness of data delivery is critical to their businesses. Today, organizations understand the importance of real-time data analysis to maintain benefits from their generated data. Real-time data analysis is also known as real-time analytics, streaming analytics, real-time streaming analytics, and event processing. Stream processing is the key to getting results in real-time. It allows us to process the data stream in real-time as it arrives. The concept of streaming data means the data are generated dynamically, and the full stream is unknown or even infinite. This data becomes …
Untapped Potential Of Clinical Text For Opioid Surveillance, Amy L. Olex, Tamas Gal, Majid Afshar, Dmitriy Dligach, Niranjan Karnik, Travis Oakes, Brihat Sharma, Meng Xie, Bridget T. Mcinnes, Julian Solway, Abel Kho, William Cramer, F. Gerard Moeller
Untapped Potential Of Clinical Text For Opioid Surveillance, Amy L. Olex, Tamas Gal, Majid Afshar, Dmitriy Dligach, Niranjan Karnik, Travis Oakes, Brihat Sharma, Meng Xie, Bridget T. Mcinnes, Julian Solway, Abel Kho, William Cramer, F. Gerard Moeller
Wright Center for Clinical and Translational Research Works
Accurate surveillance is needed to combat the growing opioid epidemic. To investigate the potential volume of missed opioid overdoses, we compare overdose encounters identified by ICD-10-CM codes and an NLP pipeline from two different medical systems. Our results show that the NLP pipeline identified a larger percentage of OOD encounters than ICD-10-CM codes. Thus, incorporating sophisticated NLP techniques into current diagnostic methods has the potential to improve surveillance on the incidence of opioid overdoses.
Radically Simplifying Gated Recurrent Architectures Without Loss Of Performance, Jonathan Boardman, Ying Xie
Radically Simplifying Gated Recurrent Architectures Without Loss Of Performance, Jonathan Boardman, Ying Xie
Published and Grey Literature from PhD Candidates
Long Short-Term Memory (LSTM) units are a family of Recurrent Neural Network (RNN) architectures that have proven incredibly effective at learning from sequence data. They are also extremely complex, making them expensive to train and difficult to understand. A recent trend towards simplification has produced the Gated Recurrent Unit (GRU) and the Minimal Gated Unit (MGU), both of which perform as well as the LSTM (or better) on a variety of tasks. The MGU is one of the simplest gated recurrent architectures at the moment. Our study demonstrates that it is possible to radically simplify the MGU without significant loss …
Canadian Hockey Leagues Game-To-Game Performance, Nick R. Riccardi
Canadian Hockey Leagues Game-To-Game Performance, Nick R. Riccardi
Sport Management - All Scholarship
This study examines game-to-game performance of players across the three Canadian Hockey Leagues (Western Hockey League, Ontario Hockey League, and Quebec Major Junior Hockey League) for the 2017-2018 season. It tests the importance of factors such as rest, travel, weather conditions, and more. Data for this study were collected from each of the three CHL websites and from www.weatherunderground.com. The null hypotheses of different factors affecting performance were tested through regression models using Ordinary Least Squares. The dependent variables, used across different specifications, were on-ice performance variables such as points, goals, and penalty minutes on a per-game basis.
Enrollment And Assessment Of A First-Year College Class Social Network For A Controlled Trial Of The Indirect Effect Of A Brief Motivational Intervention, Nancy P. Barnett, Melissa A. Clark, Shannon R. Kenney, Graham Diguiseppi, Matthew K. Meisel, Sara Balestrieri, Miles Q. Ott, John Light
Enrollment And Assessment Of A First-Year College Class Social Network For A Controlled Trial Of The Indirect Effect Of A Brief Motivational Intervention, Nancy P. Barnett, Melissa A. Clark, Shannon R. Kenney, Graham Diguiseppi, Matthew K. Meisel, Sara Balestrieri, Miles Q. Ott, John Light
Statistical and Data Sciences: Faculty Publications
Heavy drinking and its consequences among college students represent a serious public health problem, and peer social networks are a robust predictor of drinking-related risk behaviors. In a recent trial, we administered a Brief Motivational Intervention (BMI) to a small number of first-year college students to assess the indirect effects of the intervention on peers not receiving the intervention. Objectives: To present the research design, describe the methods used to successfully enroll a high proportion of a first-year college class network, and document participant characteristics. Methods: Prior to study enrollment, we consulted with a student advisory group and campus stakeholders …
Data Sharing At Scale: A Heuristic For Affirming Data Cultures, Lindsay Poirier, Brandon Costelloe-Kuehn
Data Sharing At Scale: A Heuristic For Affirming Data Cultures, Lindsay Poirier, Brandon Costelloe-Kuehn
Statistical and Data Sciences: Faculty Publications
Addressing the most pressing contemporary social, environmental, and technological challenges will require integrating insights and sharing data across disciplines, geographies, and cultures. Strengthening international data sharing networks will not only demand advancing technical, legal, and logistical infrastructure for publishing data in open, accessible formats; it will also require recognizing, respecting, and learning to work across diverse data cultures. This essay introduces a heuristic for pursuing richer characterizations of the “data cultures” at play in international, interdisciplinary data sharing. The heuristic prompts cultural analysts to query the contexts of data sharing for a particular discipline, institution, geography, or project at seven …
Classification As Catachresis: Double Binds Of Representing Difference With Semiotic Infrastructure, Lindsay Poirier
Classification As Catachresis: Double Binds Of Representing Difference With Semiotic Infrastructure, Lindsay Poirier
Statistical and Data Sciences: Faculty Publications
Background; This article explores the results of a three-year ethnographic study of how semiotic infrastructures-or digital standards and frameworks such as taxonomies, schemas, and ontologies that encode the meaning of data-are designed. Analysis: It examines debates over best practices in semiotic infrastructure design, such as how much complexity adopted languages should characterize versus how restrictive they should be. It also discusses political and pragmatic considerations that impact what and how information is represented in an information system. Conclusion and implications: This article suggests that all databased representations are forms of data power, and that examining semiotic infrastructure design provides insight …
Deep Patient Representation Of Clinical Notes Via Multi-Task Learning For Mortality Prediction, Yuqi Si, Kirk Roberts
Deep Patient Representation Of Clinical Notes Via Multi-Task Learning For Mortality Prediction, Yuqi Si, Kirk Roberts
Faculty, Staff and Student Publications
We propose a deep learning-based multi-task learning (MTL) architecture focusing on patient mortality predictions from clinical notes. The MTL framework enables the model to learn a patient representation that generalizes to a variety of clinical prediction tasks. Moreover, we demonstrate how MTL enables small but consistent gains on a single classification task (e.g., in-hospital mortality prediction) simply by incorporating related tasks (e.g., 30-day and 1-year mortality prediction) into the MTL framework. To accomplish this, we utilize a multi-level Convolutional Neural Network (CNN) associated with a MTL loss component. The model is evaluated with 3, 5, and 20 tasks and is …
The Global Disinformation Order: 2019 Global Inventory Of Organised Social Media Manipulation, Samantha Bradshaw, Philip N. Howard
The Global Disinformation Order: 2019 Global Inventory Of Organised Social Media Manipulation, Samantha Bradshaw, Philip N. Howard
Copyright, Fair Use, Scholarly Communication, etc.
Executive Summary
Over the past three years, we have monitored the global organization of social media manipulation by governments and political parties. Our 2019 report analyses the trends of computational propaganda and the evolving tools, capacities, strategies, and resources.
1. Evidence of organized social media manipulation campaigns which have taken place in 70 countries, up from 48 countries in 2018 and 28 countries in 2017. In each country, there is at least one political party or government agency using social media to shape public attitudes domestically.
2.Social media has become co-opted by many authoritarian regimes. In 26 countries, computational propaganda …
Detecting Special-Cause Variation 'Events' From Process Data Signatures, Timothy M. Young, Olga Khaliukova, Nicolas André, Alexander Petutschnigg, Timothy G. Rials, Chung-Hao Chen
Detecting Special-Cause Variation 'Events' From Process Data Signatures, Timothy M. Young, Olga Khaliukova, Nicolas André, Alexander Petutschnigg, Timothy G. Rials, Chung-Hao Chen
Electrical & Computer Engineering Faculty Publications
The ability to detect the special-cause variation of incoming feedstocks from advanced sensor technology is invaluable to manufacturers. Many on-line sensors produce data signatures that require further off-line statistical processing for interpretation by operational personnel. However, early detection of changes in variation in incoming feedstocks may be imperative to promote early-stage preventive measures. A method is proposed in this applied study for developing control bands to quantify the variation of data signatures in the context of statistical process control (SPC). Control bands based on pointwise prediction intervals constructed from the Bonferroni Inequality and Bayesian smoothing splines are developed. Applications using …
Population Data Centre Profile - The Western Australian Data Linkage Branch, Steve Hodges, Tom Eitelhuber, Alexandra Merchant, Janine Alan
Population Data Centre Profile - The Western Australian Data Linkage Branch, Steve Hodges, Tom Eitelhuber, Alexandra Merchant, Janine Alan
Research outputs 2014 to 2021
Established in 1995, the Western Australian Data Linkage Branch (DLB) is Australia’s longest running data linkage agency. The Western Australian Data Linkage System (WADLS) employs an enduring linkage model spanning over 60 data collections supported by internally developed and supported software and IT infrastructure. DLB has delivered, and continues to deliver, a range of significant data linkage innovations, many of which have been adopted elsewhere. A current restructure within the Western Australian Department of Health (which we will refer to as the Department of Health) will provide an improved funding model geared toward addressing issues with staff retention, capacity and …
Explorobot: Rapid Exploration With Chart Automation, Tamara Matthews, Rohan Goel, John Mcauley
Explorobot: Rapid Exploration With Chart Automation, Tamara Matthews, Rohan Goel, John Mcauley
Conference papers
General-purpose visualization tools are used by people with varying degrees of data literacy. Often the user is not a professional analyst or data scientist and uses the tool infrequently, to support an aspect of their job. This can present difficulties as the user’s unfamiliarity with visualization practice and infrequent use of the tool can result in long processing time, inaccurate data representations or inappropriate visual encodings. To address this problem, we developed a visual analytics application called exploroBOT. The exploroBOT automatically generates visualizations and the exploration guidance path (an associated network of decision points, mapping nodes where visualizations change). These …
On The Inability Of Markov Models To Capture Criticality In Human Mobility, Vaibhav Klukarni, Abhijit Mahalunkar, Benoit Garbinato, John Kelleher
On The Inability Of Markov Models To Capture Criticality In Human Mobility, Vaibhav Klukarni, Abhijit Mahalunkar, Benoit Garbinato, John Kelleher
Conference papers
We examine the non-Markovian nature of human mobility by exposing the inability of Markov models to capture criticality in human mobility. In particular, the assumed Markovian nature of mobility was used to establish an upper bound on the predictability of human mobility, based on the temporal entropy. Since its inception, this bound has been widely used for validating the performance of mobility prediction models. We show that the variants of recurrent neural network architectures can achieve significantly higher prediction accuracy surpassing this upper bound. The central objective of our work is to show that human-mobility dynamics exhibit criticality characteristics which …
Knowledge Management Overview Of Feature Selection Problem In High-Dimensional Financial Data: Cooperative Co-Evolution And Map Reduce Perspectives, A. N. M. Bazlur Rashid, Tonmoy Choudhury
Knowledge Management Overview Of Feature Selection Problem In High-Dimensional Financial Data: Cooperative Co-Evolution And Map Reduce Perspectives, A. N. M. Bazlur Rashid, Tonmoy Choudhury
Research outputs 2014 to 2021
The term "big data" characterizes the massive amounts of data generation by the advanced technologies in different domains using 4Vs volume, velocity, variety, and veracity-to indicate the amount of data that can only be processed via computationally intensive analysis, the speed of their creation, the different types of data, and their accuracy. High-dimensional financial data, such as time-series and space-Time data, contain a large number of features (variables) while having a small number of samples, which are used to measure various real-Time business situations for financial organizations. Such datasets are normally noisy, and complex correlations may exist between their features, …
Profiling And Identifying Individual Usersby Their Command Line Usage And Writing Style, Darusalam Darusalam, Helen Ashman
Profiling And Identifying Individual Usersby Their Command Line Usage And Writing Style, Darusalam Darusalam, Helen Ashman
Knowledge Engineering and Data Science
Profiling and identifying individual users is an approach for intrusion detection in a computer system. User profiles are important in many applications since they record highly user-specific information - profiles are basically built to record information about users or for users to share experiences with each other. This research extends previous research on re-authenticating users with their user profiles. This research focuses on the potential to add psychometric user characteristics into the user model so as to be able to detect unauthorized users who may be masquerading as a genuine user. There are five participants involved in the investigation for …
Energy Efficiency Metrics Of University Data Centers, Leonel Hernandez, Genett Jimenez, Piedad Marchena
Energy Efficiency Metrics Of University Data Centers, Leonel Hernandez, Genett Jimenez, Piedad Marchena
Knowledge Engineering and Data Science
The data centers are fundamental pieces in the network and computing infrastructure,and evidently today more than ever they are relevant. Since they support the processing, analysis, assurance of the data generated in the network and by the applications in the cloud, which every day increases its volume thanks to technologies such as Internet of Things, Virtualization, and cloud computing, among others. Precisely the management of this large volume of information makes the data centers consume a lot of energy, generating great concern to owners and administrators. Green Data Centers offer a solution to this problem, reducing the impact produced by …
Signature Pattern Recognition Using Kohonen Network, Nadia Roosmalita Sari, Mohammad Zoqi Sarwani, Yudha Alif Aulia, Wayan Firdaus Mahmudy
Signature Pattern Recognition Using Kohonen Network, Nadia Roosmalita Sari, Mohammad Zoqi Sarwani, Yudha Alif Aulia, Wayan Firdaus Mahmudy
Knowledge Engineering and Data Science
A signature is a special form of handwriting that used for human identification process. The current identification process is extremely ineffective. People have to manually compare signatures with the previously stored data. This study proposed SOM Kohonen algorithm as the method of signature pattern recognition. This method has able to visualize high-dimensional data. The image processing method is used in this study in pre-processing data phase. The accuracy of SOM Kohonen was 70 %, indicated the method used was good enough for pattern recognition.