Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 3901 - 3930 of 7251

Full-Text Articles in Databases and Information Systems

An Adaptive Computational Model For Personalized Persuasion, Yilin Kang, Ah-Hwee Tan, Chunyan Miao Jul 2015

An Adaptive Computational Model For Personalized Persuasion, Yilin Kang, Ah-Hwee Tan, Chunyan Miao

Research Collection School Of Computing and Information Systems

While a variety of persuasion agents have been created and applied in different domains such as marketing, military training and health industry, there is a lack of a model which can provide a unified framework for different persuasion strategies. Specifically, persuasion is not adaptable to the individuals’ personal states in different situations. Grounded in the Elaboration Likelihood Model (ELM), this paper presents a computational model called Model for Adaptive Persuasion (MAP) for virtual agents. MAP is a semi-connected network model which enables an agent to adapt its persuasion strategies through feedback. We have implemented and evaluated a MAP-based virtual nurse …


Landmark Classification With Hierarchical Multi-Modal Exemplar Feature, Lei Zhu, Jialie Shen, Hai Jin, Liang Xie, Ran Zheng Jul 2015

Landmark Classification With Hierarchical Multi-Modal Exemplar Feature, Lei Zhu, Jialie Shen, Hai Jin, Liang Xie, Ran Zheng

Research Collection School Of Computing and Information Systems

Landmark image classification attracts increasing research attention due to its great importance in real applications, ranging from travel guide recommendation to 3-D modelling and visualization of geolocation. While large amount of efforts have been invested, it still remains unsolved by academia and industry. One of the key reasons is the large intra-class variance rooted from the diverse visual appearance of landmark images. Distinguished from most existing methods based on scalable image search, we approach the problem from a new perspective and model landmark classification as multi-modal categorization, which enjoys advantages of low storage overhead and high classification efficiency. Toward this …


State Preserving Extreme Learning Machine For Face Recognition, Md. Zahangir Alom, Paheding Sidike, Vijayan K. Asari, Tarek M. Taha Jul 2015

State Preserving Extreme Learning Machine For Face Recognition, Md. Zahangir Alom, Paheding Sidike, Vijayan K. Asari, Tarek M. Taha

Electrical and Computer Engineering Faculty Publications

Extreme Learning Machine (ELM) has been introduced as a new algorithm for training single hidden layer feed-forward neural networks (SLFNs) instead of the classical gradient-based algorithms. Based on the consistency property of data, which enforce similar samples to share similar properties, ELM is a biologically inspired learning algorithm with SLFNs that learns much faster with good generalization and performs well in classification applications. However, the random generation of the weight matrix in current ELM based techniques leads to the possibility of unstable outputs in the learning and testing phases. Therefore, we present a novel approach for computing the weight matrix …


Automatic Video Self Modeling For Voice Disorder, Ju Shen, Changpeng Ti, Anusha Raghunathan, Sen-Ching S. Cheung, Rita Patel Jul 2015

Automatic Video Self Modeling For Voice Disorder, Ju Shen, Changpeng Ti, Anusha Raghunathan, Sen-Ching S. Cheung, Rita Patel

Computer Science Faculty Publications

Video self modeling (VSM) is a behavioral intervention technique in which a learner models a target behavior by watching a video of him- or herself. In the field of speech language pathology, the approach of VSM has been successfully used for treatment of language in children with Autism and in individuals with fluency disorder of stuttering. Technical challenges remain in creating VSM contents that depict previously unseen behaviors. In this paper, we propose a novel system that synthesizes new video sequences for VSM treatment of patients with voice disorders. Starting with a video recording of a voice-disorder patient, the proposed …


Domain Specific Document Retrieval Framework For Real-Time Social Health Data, Swapnil Soni Jul 2015

Domain Specific Document Retrieval Framework For Real-Time Social Health Data, Swapnil Soni

Kno.e.sis Publications

With the advent of the web search and microblogging, the percentage of Online Health Information Seekers (OHIS) using these online services to share and seek health real-time information has in- creased exponentially. OHIS use web search engines or microblogging search services to seek out latest, relevant as well as reliable health in- formation. When OHIS turn to microblogging search services to search real-time content, trends and breaking news, etc. the search results are not promising. Two major challenges exist in the current microblogging search engines are keyword based techniques and results do not contain real-time information. To address these challenges, …


A Comparative Study Between Motivated Learning And Reinforcement Learning, James T. Graham, Janusz A. Starzyk, Zhen Ni, Haibo He, T.-H. Teng, Ah-Hwee Tan Jul 2015

A Comparative Study Between Motivated Learning And Reinforcement Learning, James T. Graham, Janusz A. Starzyk, Zhen Ni, Haibo He, T.-H. Teng, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

This paper analyzes advanced reinforcement learning techniques and compares some of them to motivated learning. Motivated learning is briefly discussed indicating its relation to reinforcement learning. A black box scenario for comparative analysis of learning efficiency in autonomous agents is developed and described. This is used to analyze selected algorithms. Reported results demonstrate that in the selected category of problems, motivated learning outperformed all reinforcement learning algorithms we compared with.


Fast Optimal Aggregate Point Search For A Merged Set On Road Networks, Weiwei Sun, Chong Chen, Baihua Zheng, Chunan Chen, Liang Zhu, Weimo Liu, Yan Huang Jul 2015

Fast Optimal Aggregate Point Search For A Merged Set On Road Networks, Weiwei Sun, Chong Chen, Baihua Zheng, Chunan Chen, Liang Zhu, Weimo Liu, Yan Huang

Research Collection School Of Computing and Information Systems

Aggregate nearest neighbor query, which returns an optimal target point that minimizes the aggregate distance for a given query point set, is one of the most important operations in spatial databases and their application domains. This paper addresses the problem of finding the aggregate nearest neighbor for a merged set that consists of the given query point set and multiple points needed to be selected from a candidate set, which we name as merged aggregate nearest neighbor(MANN) query. This paper proposes two algorithms to process MANN query on road networks when aggregate function is max. Then, we extend the algorithms …


A Convolution Kernel Approach To Identifying Comparisons In Text, Maksim Tkachenko, Hady W. Lauw Jul 2015

A Convolution Kernel Approach To Identifying Comparisons In Text, Maksim Tkachenko, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Comparisons in text, such as in online reviews, serve as useful decision aids. In this paper, we focus on the task of identifying whether a comparison exists between a specific pair of entity mentions in a sentence. This formulation is transformative, as previous work only seeks to determine whether a sentence is comparative, which is presumptuous in the event the sentence mentions multiple entities and is comparing only some, not all, of them. Our approach leverages not only lexical features such as salient words, but also structural features expressing the relationships among words and entity mentions. To model these features …


Online Learning To Rank For Content-Based Image Retrieval, Ji Wan, Pengcheng Wu, Steven C. H. Hoi, Peilin Zhao, Xingyu Gao, Dayong Wang, Yongdong. Zhang, Jintao Li Jul 2015

Online Learning To Rank For Content-Based Image Retrieval, Ji Wan, Pengcheng Wu, Steven C. H. Hoi, Peilin Zhao, Xingyu Gao, Dayong Wang, Yongdong. Zhang, Jintao Li

Research Collection School Of Computing and Information Systems

A major challenge in Content-Based Image Retrieval (CBIR) is to bridge the semantic gap between low-level image contents and high-level semantic concepts. Although researchers have investigated a variety of retrieval techniques using different types of features and distance functions, no single best retrieval solution can fully tackle this challenge. In a real-world CBIR task, it is often highly desired to combine multiple types of different feature representations and diverse distance measures in order to close the semantic gap. In this paper, we investigate a new framework of learning to rank for CBIR, which aims to seek the optimal combination of …


Solar: Scalable Online Learning Algorithms For Ranking, Jialei Wang, Ji Wan, Yongdong Zhang, Steven C. H. Hoi Jul 2015

Solar: Scalable Online Learning Algorithms For Ranking, Jialei Wang, Ji Wan, Yongdong Zhang, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Traditional learning to rank methods learn ranking models from training data in a batch and offline learning mode, which suffers from some critical limitations, e.g., poor scalability as the model has to be retrained from scratch whenever new training data arrives. This is clearly nonscalable for many real applications in practice where training data often arrives sequentially and frequently. To overcome the limitations, this paper presents SOLAR- a new framework of Scalable Online Learning Algorithms for Ranking, to tackle the challenge of scalable learning to rank. Specifically, we propose two novel SOLAR algorithms and analyze their IR measure bounds theoretically. …


Structured Learning From Heterogeneous Behavior For Social Identity Linkage, Siyuan Liu, Shuhui Wang, Feida Zhu Jul 2015

Structured Learning From Heterogeneous Behavior For Social Identity Linkage, Siyuan Liu, Shuhui Wang, Feida Zhu

Research Collection School Of Computing and Information Systems

Social identity linkage across different social media platforms is of critical importance to business intelligence by gaining from social data a deeper understanding and more accurate profiling of users. In this paper, we propose a solution framework, HYDRA, which consists of three key steps: (I) we model heterogeneous behavior by long-term topical distribution analysis and multi-resolution temporal behavior matching against high noise and information missing, and the behavior similarity are described by multi-dimensional similarity vector for each user pair; (II) we build structure consistency models to maximize the structure and behavior consistency on users' core social structure across different platforms, …


Using Tweets To Help Sentence Compression For News Highlights Generation, Zhongyu Wei, Yang Liu, Chen Li, Wei Gao Jul 2015

Using Tweets To Help Sentence Compression For News Highlights Generation, Zhongyu Wei, Yang Liu, Chen Li, Wei Gao

Research Collection School Of Computing and Information Systems

We explore using relevant tweets of a given news article to help sentence compression for generating compressive news highlights. We extend an unsupervised dependency-tree based sentence compression approach by incorporating tweet information to weight the tree edge in terms of informativeness and syntactic importance. The experimental results on a public corpus that contains both news articles and relevant tweets show that our proposed tweets guided sentence compression method can improve the summarization performance significantly compared to the baseline generic sentence compression method.


Personalized Sentiment Classification Based On Latent Individuality Of Microblog Users, Kaisong Song, Shi Feng, Wei Gao, Daling Wang, Ge Yu, Kam-Fai Wong Jul 2015

Personalized Sentiment Classification Based On Latent Individuality Of Microblog Users, Kaisong Song, Shi Feng, Wei Gao, Daling Wang, Ge Yu, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Sentiment expression in microblog posts often reflects user’s specific individuality due to different language habit, personal character, opinion bias and so on. Existing sentiment classification algorithms largely ignore such latent personal distinctions among different microblog users. Meanwhile, sentiment data of microblogs are sparse for individual users, making it infeasible to learn effective personalized classifier. In this paper, we propose a novel, extensible personalized sentiment classification method based on a variant of latent factor model to capture personal sentiment variations by mapping users and posts into a low-dimensional factor space. We alleviate the sparsity of personal texts by decomposing the posts …


A Hassle-Free Unsupervised Domain Adaptation Method Using Instance Similarity Features, Jianfei Yu, Jing Jiang Jul 2015

A Hassle-Free Unsupervised Domain Adaptation Method Using Instance Similarity Features, Jianfei Yu, Jing Jiang

Research Collection School Of Computing and Information Systems

We present a simple yet effective unsupervised domain adaptation method that can be generally applied for different NLP tasks. Our method uses unlabeled target domain instances to induce a set of instance similarity features. These features are then combined with the original features to represent labeled source domain instances. Using three NLP tasks, we show that our method consistently out-performs a few baselines, including SCL, an existing general unsupervised domain adaptation method widely used in NLP. More importantly, our method is very easy to implement and incurs much less computational cost than SCL.


Log-Euclidean Metric Learning On Symmetric Positive Definite Manifold With Application To Image Set Classification, Zhiwu Huang, R. Wang, S. Shan, X. Li, X. Chen Jul 2015

Log-Euclidean Metric Learning On Symmetric Positive Definite Manifold With Application To Image Set Classification, Zhiwu Huang, R. Wang, S. Shan, X. Li, X. Chen

Research Collection School Of Computing and Information Systems

The manifold of Symmetric Positive Definite (SPD) matrices has been successfully used for data representation in image set classification. By endowing the SPD manifold with Log-Euclidean Metric, existing methods typically work on vector-forms of SPD matrix logarithms. This however not only inevitably distorts the geometrical structure of the space of SPD matrix logarithms but also brings low efficiency especially when the dimensionality of SPD matrix is high. To overcome this limitation, we propose a novel metric learning approach to work directly on logarithms of SPD matrices. Specifically, our method aims to learn a tangent map that can directly transform the …


Mobile Phishing Attacks And Mitigation Techniques, Hossain Shahriar, Tulin Klintic, Victor Clincy Jun 2015

Mobile Phishing Attacks And Mitigation Techniques, Hossain Shahriar, Tulin Klintic, Victor Clincy

Faculty Articles

Mobile devices have taken an essential role in the portable computer world. Portability, small screen size, and lower cost of production make these devices popular replacements for desktop and laptop computers for many daily tasks, such as surfing on the Internet, playing games, and shopping online. The popularity of mobile devices such as tablets and smart phones has made them a frequent target of traditional web-based attacks, especially phishing. Mobile device-based phishing takes its share of the pie to trick users into entering their credentials in fake websites or fake mobile applications. This paper discusses various phishing attacks using mobile …


Qcri: Answer Selection For Community Question Answering - Experiment For Arabic And English, Massimo Nicosia, Simone Filice, Alberto Barron-Cedeno, Iman Saleh, Hamdy Mubarak, Wei Gao, Preslav Nakov, Giovanni Da San Martino, Alessandro Moschitti, Kareem Darwish, Lluis Marquz Marquz, Shafiq Joty, Walid Magdy Magdy Jun 2015

Qcri: Answer Selection For Community Question Answering - Experiment For Arabic And English, Massimo Nicosia, Simone Filice, Alberto Barron-Cedeno, Iman Saleh, Hamdy Mubarak, Wei Gao, Preslav Nakov, Giovanni Da San Martino, Alessandro Moschitti, Kareem Darwish, Lluis Marquz Marquz, Shafiq Joty, Walid Magdy Magdy

Research Collection School Of Computing and Information Systems

This paper describes QCRI’s participation in SemEval-2015 Task 3 “Answer Selection in Community Question Answering”, which targeted real-life Web forums, and was offered in both Arabic and English. We apply a supervised machine learning approach considering a manifold of features including among others word n-grams, text similarity, sentiment analysis, the presence of specific words, and the context of a comment. Our approach was the best performing one in the Arabic subtask and the third best in the two English subtasks


"Time For Dabs": Analyzing Twitter Data On Butane Hash Oil Use, Raminta Daniulaityte, Robert G. Carlson, Farahnaz Golroo, Sanjaya Wijeratne, Edward W. Boyer, Silvia S. Martins, Ramzi W. Nahhas, Amit P. Sheth Jun 2015

"Time For Dabs": Analyzing Twitter Data On Butane Hash Oil Use, Raminta Daniulaityte, Robert G. Carlson, Farahnaz Golroo, Sanjaya Wijeratne, Edward W. Boyer, Silvia S. Martins, Ramzi W. Nahhas, Amit P. Sheth

Kno.e.sis Publications

No abstract provided.


Trust Management: Multimodal Data Perspective, Krishnaprasad Thirunarayan Jun 2015

Trust Management: Multimodal Data Perspective, Krishnaprasad Thirunarayan

Kno.e.sis Publications

No abstract provided.


Dynamic Redeployment To Counter Congestion Or Starvation In Vehicle Sharing Systems, Supriyo Ghosh, Pradeep Varakantham, Yossiri Adulyasak, Patrick Jaillet Jun 2015

Dynamic Redeployment To Counter Congestion Or Starvation In Vehicle Sharing Systems, Supriyo Ghosh, Pradeep Varakantham, Yossiri Adulyasak, Patrick Jaillet

Research Collection School Of Computing and Information Systems

Extensive usage of private vehicles has led to increased traffic congestion, carbon emissions, and usage of non-renewable resources. These concerns have led to the wide adoption of vehicle sharing (ex: bike sharing, car sharing) systems in many cities of the world. In vehicle-sharing systems, base stations (ex: docking stations for bikes) are strategically placed throughout a city and each of the base stations contain a pre-determined number of vehicles at the beginning of each day. Due to the stochastic and individualistic movement of customers,there is typically either congestion (more than required)or starvation (fewer than required) of vehicles at certain base …


A Modular Approach For Key-Frame Selection In Wide Area Surveillance Video Analysis, Almabrok Essa, Paheding Sidike, Vijayan K. Asari Jun 2015

A Modular Approach For Key-Frame Selection In Wide Area Surveillance Video Analysis, Almabrok Essa, Paheding Sidike, Vijayan K. Asari

Electrical and Computer Engineering Faculty Publications

This paper presents an efficient preprocessing algorithm for big data analysis. Our proposed key-frame selection method utilizes the statistical differences among subsequent frames to automatically select only the frames that contain the desired contextual information and discard the rest of the insignificant frames.

We anticipate that such key frame selection technique will have significant impact on wide area surveillance applications such as automatic object detection and recognition in aerial imagery. Three real-world datasets are used for evaluation and testing and the observed results are encouraging.


Geospatial Data Modeling To Support Energy Pipeline Integrity Management, Austin Wylie Jun 2015

Geospatial Data Modeling To Support Energy Pipeline Integrity Management, Austin Wylie

Master's Theses

Several hundred thousand miles of energy pipelines span the whole of North America -- responsible for carrying the natural gas and liquid petroleum that power the continent's homes and economies. These pipelines, so crucial to everyday goings-on, are closely monitored by various operating companies to ensure they perform safely and smoothly.

Happenings like earthquakes, erosion, and extreme weather, however -- and human factors like vehicle traffic and construction -- all pose threats to pipeline integrity. As such, there is a tremendous need to measure and indicate useful, actionable data for each region of interest, and operators often use computer-based decision …


An Examination Of Service Level Agreement Attributes That Influence Cloud Computing Adoption, Howard Gregory Hamilton Jun 2015

An Examination Of Service Level Agreement Attributes That Influence Cloud Computing Adoption, Howard Gregory Hamilton

CCAC Theses and Dissertations

Cloud computing is perceived as the technological innovation that will transform future investments in information technology. As cloud services become more ubiquitous, public and private enterprises still grapple with concerns about cloud computing. One such concern is about service level agreements (SLAs) and their appropriateness.

While the benefits of using cloud services are well defined, the debate about the challenges that may inhibit the seamless adoption of these services still continues. SLAs are seen as an instrument to help foster adoption. However, cloud computing SLAs are alleged to be ineffective, meaningless, and costly to administer. This could impact widespread acceptance …


Service Quality And Perceived Value Of Cloud Computing-Based Service Encounters: Evaluation Of Instructor Perceived Service Quality In Higher Education In Texas, Eges Egedigwe Jun 2015

Service Quality And Perceived Value Of Cloud Computing-Based Service Encounters: Evaluation Of Instructor Perceived Service Quality In Higher Education In Texas, Eges Egedigwe

CCAC Theses and Dissertations

Cloud computing based technology is becoming increasingly popular as a way to deliver quality education to community colleges, universities and other organizations. At the same time, compared with other industries, colleges have been slow on implementing and sustaining cloud computing services on an institutional level because of budget constraints facing many large community colleges, in addition to other obstacles. Faced with this challenge, key stakeholders are increasingly realizing the need to focus on service quality as a measure to improve their competitive position in today's highly competitive environment. Considering the amount of study done with cloud computing in education, very …


The Evolution Of Scientific Productivity Of Junior Scholars, Chun-Hua Tsai, Yu-Ru Lin Jun 2015

The Evolution Of Scientific Productivity Of Junior Scholars, Chun-Hua Tsai, Yu-Ru Lin

Information Systems and Quantitative Analysis Faculty Proceedings & Presentations

Publishing academic work has been recognized as a key indicator for measuring scholars’ scientific productivity and having crucial impact on their future career. However, little has been known about how the majority of researchers progress in publishing papers across disciplines. In this work, using a collection consisting of over five millions academic publications across 15 disciplines, we study how the scientific productivity patterns of junior scholars change across different generations and different domains. Our study results help understand the evolution of the competitive “publish or perish” academic culture.


Dietary Microrna Database (Dmd): An Archive Database And Analytic Tool For Food-Borne Micrornas, Kevin Chiang, Jiang Shu, Janos Zempleni, Juan Cui Jun 2015

Dietary Microrna Database (Dmd): An Archive Database And Analytic Tool For Food-Borne Micrornas, Kevin Chiang, Jiang Shu, Janos Zempleni, Juan Cui

School of Computing: Faculty Publications

With the advent of high throughput technology, a huge amount of microRNA information has been added to the growing body of knowledge for non-coding RNAs. Here we present the Dietary MicroRNA Databases (DMD), the first repository for archiving and analyzing the published and novel microRNAs discovered in dietary resources. Currently there are fifteen types of dietary species, such as apple, grape, cow milk, and cow fat, included in the database originating from 9 plant and 5 animal species. Annotation for each entry, a mature microRNA indexed as DM0000*, covers information of the mature sequences, genome locations, hairpin structures of parental …


Semi-Supervised Domain Adaptation With Subspace Learning For Visual Recognition, Ting Yao, Yingwei Pan, Chong-Wah Ngo, Houqiang Li, Tao Mei Jun 2015

Semi-Supervised Domain Adaptation With Subspace Learning For Visual Recognition, Ting Yao, Yingwei Pan, Chong-Wah Ngo, Houqiang Li, Tao Mei

Research Collection School Of Computing and Information Systems

In many real-world applications, we are often facing the problem of cross domain learning, i.e., to borrow the labeled data or transfer the already learnt knowledge from a source domain to a target domain. However, simply applying existing source data or knowledge may even hurt the performance, especially when the data distribution in the source and target domain is quite different, or there are very few labeled data available in the target domain. This paper proposes a novel domain adaptation framework, named Semi-supervised Domain Adaptation with Subspace Learning (SDASL), which jointly explores invariant lowdimensional structures across domains to correct data …


Emif: Towards A Scalable And Effective Indexing Framework For Large Scale Music Retrieval, Jialie Shen, Tao Mei, Dacheng Tao, Xuelong Li, Yong Rui Jun 2015

Emif: Towards A Scalable And Effective Indexing Framework For Large Scale Music Retrieval, Jialie Shen, Tao Mei, Dacheng Tao, Xuelong Li, Yong Rui

Research Collection School Of Computing and Information Systems

This article presents a novel indexing framework called EMIF (Effective Music Indexing Framework) to facilitate scalable and accurate content based music retrieval. EMIF system architecture is designed based on a "classification-and-indexing" principle and consists of two main functionality layers: 1) a novel semantic-sensitive classification to identify input music's category and 2) multiple indexing structures - one local indexing structure corresponds to one semantic category. EMIF's layered architecture not only enables superior search accuracy but also reduces query response time significantly. To evaluate the system, a set of comprehensive experimental studies have been carried out using large test collection and EMIF …


Face Video Retrieval With Image Query Via Hashing Across Euclidean Space And Riemannian Manifold, Y. Li, R. Wang, Zhiwu Huang, S. Shan, X. Chen Jun 2015

Face Video Retrieval With Image Query Via Hashing Across Euclidean Space And Riemannian Manifold, Y. Li, R. Wang, Zhiwu Huang, S. Shan, X. Chen

Research Collection School Of Computing and Information Systems

Retrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video clip are usually represented quite differently. Typically, face image is represented as point (i.e., vector) in Euclidean space, while video clip is seemingly modeled as a point (e.g., covariance matrix) on some particular Riemannian manifold in the light of its recent promising success. It thus incurs a new hashing-based retrieval problem of matching …


Should We Use The Sample? Analyzing Datasets Sampled From Twitter's Stream Api, Yazhe Wang, Jamie Callan, Baihua Zheng Jun 2015

Should We Use The Sample? Analyzing Datasets Sampled From Twitter's Stream Api, Yazhe Wang, Jamie Callan, Baihua Zheng

Research Collection School Of Computing and Information Systems

Researchers have begun studying content obtained from microblogging services such as Twitter to address a variety of technological, social, and commercial research questions. The large number of Twitter users and even larger volume of tweets often make it impractical to collect and maintain a complete record of activity; therefore, most research and some commercial software applications rely on samples, often relatively small samples, of Twitter data. For the most part, sample sizes have been based on availability and practical considerations. Relatively little attention has been paid to how well these samples represent the underlying stream of Twitter data. To fill …