Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2017

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 241 - 270 of 373

Full-Text Articles in Databases and Information Systems

An Evidence-Based Review Of Academic Web Search Engines, 2014-2016: Implications For Librarians’ Practice And Research Agenda, Jody C. Fagan Mar 2017

An Evidence-Based Review Of Academic Web Search Engines, 2014-2016: Implications For Librarians’ Practice And Research Agenda, Jody C. Fagan

Libraries

Academic web search engines have become central to scholarly research. While the fitness of Google Scholar for research purposes has been examined repeatedly, Microsoft Academic and Google Books have not received much attention. Recent studies have much to tell us about the coverage and utility of Google Scholar, its coverage of the sciences, and its utility for evaluating researcher impact. But other aspects have been understudied, such as coverage of the arts and humanities, books, and non-Western, non-English publications. User research has also tapered off. A small number of articles hint at the opportunity for librarians to become expert advisors …


Ten Simple Rules For Responsible Big Data Research, Matthew Zook, Solon Barocas, Danah Boyd, Kate Crawford, Emily Keller, Seeta Peña Gangadharan, Alyssa Goodman, Rachelle Hollander, Barbara A. Koenig, Jacob Metcalf, Arvind Narayanan, Alondra Nelson, Frank Pasquale Mar 2017

Ten Simple Rules For Responsible Big Data Research, Matthew Zook, Solon Barocas, Danah Boyd, Kate Crawford, Emily Keller, Seeta Peña Gangadharan, Alyssa Goodman, Rachelle Hollander, Barbara A. Koenig, Jacob Metcalf, Arvind Narayanan, Alondra Nelson, Frank Pasquale

Geography Faculty Publications

No abstract provided.


Forging Blockchains: Spatial Production And Political Economy Of Decentralized Cryptocurrency Code/Spaces, Joe Blankenship Mar 2017

Forging Blockchains: Spatial Production And Political Economy Of Decentralized Cryptocurrency Code/Spaces, Joe Blankenship

USF Tampa Graduate Theses and Dissertations

Cryptocurrencies and blockchains are increasingly used, implemented and adapted for numerous purposes; people and businesses are integrating these technologies into their practices and strategies, creating new political economies and spaces in and of everyday life. This thesis seeks to develop a foundation of geographic theory for the study of spatial production within and surrounding blockchain technologies focusing on acute studies of Bitcoin as cryptocurrency, Ethereum as digital marketplace, and their conditions of possibility as decentralized autonomous organizations. Utilizing concepts from Henri Lefebvre's Production of Space, this thesis situates blockchain technologies within the wider discussion about the political economy of …


Iota Pi Application For Ios, Deborah Newberry Mar 2017

Iota Pi Application For Ios, Deborah Newberry

Computer Science and Software Engineering

Kappa Kappa Psi is a national honorary fraternity for college band members. They meet every Sunday night, and during these meetings they plan events (both internal and external) that aim to meet our goal of making sure that the Cal Poly band programs have their social, financial, material, and educational needs satisfied. This paper details the steps I took to create an application for them to use internally to help ease organizational processes.


Improving Automated Bug Triaging With Specialized Topic Model, Xin Xia, David Lo, Ying Ding, Jafar M. Al-Kofahi, Tien N. Nguyen, Xinyu Wang Mar 2017

Improving Automated Bug Triaging With Specialized Topic Model, Xin Xia, David Lo, Ying Ding, Jafar M. Al-Kofahi, Tien N. Nguyen, Xinyu Wang

Research Collection School Of Computing and Information Systems

Bug triaging refers to the process of assigning a bug to the most appropriate developer to fix. It becomes more and more difficult and complicated as the size of software and the number of developers increase. In this paper, we propose a new framework for bug triaging, which maps the words in the bug reports (i.e., the term space) to their corresponding topics (i.e., the topic space). We propose a specialized topic modeling algorithm named multi-feature topic model (MTM) which extends Latent Dirichlet Allocation (LDA) for bug triaging. MTM considers product and component information of bug reports to map the …


Inferring User Consumption Preferences From Social Media, Yang Li, Jing Jiang, Ting Liu Mar 2017

Inferring User Consumption Preferences From Social Media, Yang Li, Jing Jiang, Ting Liu

Research Collection School Of Computing and Information Systems

Social Media has already become a new arena of our lives and involved different aspects of our social presence. Users' personal information and activities on social media presumably reveal their personal interests, which offer great opportunities for many e-commerce applications. In this paper, we propose a principled latent variable model to infer user consumption preferences at the category level (e.g. inferring what categories of products a user would like to buy). Our model naturally links users' published content and following relations on microblogs with their consumption behaviors on e-commerce websites. Experimental results show our model outperforms the state-of-the-art methods significantly …


Metric Similarity Joins Using Mapreduce, Yunjun Gao, Keyu Yang, Lu Chen, Baihua Zheng, Gang Chen, Chun Chen Mar 2017

Metric Similarity Joins Using Mapreduce, Yunjun Gao, Keyu Yang, Lu Chen, Baihua Zheng, Gang Chen, Chun Chen

Research Collection School Of Computing and Information Systems

Given two object sets Q and O , a metric similarity join finds similar object pairs according to a certain criterion. This operation has a wide variety of applications in data cleaning, data mining, to name but a few. However, the rapidly growing volume of data nowadays challenges traditional metric similarity join methods, and thus, a distributed method is required. In this paper, we adopt a popular distributed framework, namely, MapReduce, to support scalable metric similarity joins. To ensure the load balancing, we present two sampling based partition methods. One utilizes the pivot and the space-filling curve mappings to cluster …


Effective K-Vertex Connected Component Detection In Large-Scale Networks, Yuan Li, Yuha Zhao, Guoren Wang, Feida Zhu, Yubao Wu, Shenglei Shi Mar 2017

Effective K-Vertex Connected Component Detection In Large-Scale Networks, Yuan Li, Yuha Zhao, Guoren Wang, Feida Zhu, Yubao Wu, Shenglei Shi

Research Collection School Of Computing and Information Systems

Finding components with high connectivity is an important problem in component detection with a wide range of applications, e.g., social network analysis, web-page research and bioinformatics. In particular, k-edge connected component (k-ECC) has recently been extensively studied to discover disjoint components. Yet many real applications present needs and challenges for overlapping components. In this paper, we propose a k-vertex connected component (k-VCC) model, which is much more cohesive and therefore allows overlapping between components. To find k-VCCs, a top-down framework is first developed to find the exact k-VCCs. To further reduce the high computational cost for input networks of large …


Version-Sensitive Mobile App Recommendation, Da Cao, Liqiang Nie, Xiangnan He, Xiaochi Wei, Jialie Shen, Shunxiang Wu, Tat-Seng Chua Mar 2017

Version-Sensitive Mobile App Recommendation, Da Cao, Liqiang Nie, Xiangnan He, Xiaochi Wei, Jialie Shen, Shunxiang Wu, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Being part and parcel of the daily life for billions of people all over the globe, the domain of mobile Applications (Apps) is the fastest growing sector of mobile market today. Users, however, are frequently overwhelmed by the vast number of released Apps and frequently updated versions. Towards this end, we propose a novel version-sensitive mobile App recommendation framework. It is able to recommend appropriate Apps to right users by jointly exploring the version progression and dual-heterogeneous data. It is helpful for alleviating the data sparsity problem caused by version division. As a byproduct, it can be utilized to solve …


Social Tag Relevance Learning Via Ranking-Oriented Neighbor Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian Mar 2017

Social Tag Relevance Learning Via Ranking-Oriented Neighbor Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian

Research Collection School Of Computing and Information Systems

High quality tags play a critical role in applications involving online multimedia search, such as social image annotation, sharing and browsing. However, user-generated tags in real world are often imprecise and incomplete to describe the image contents, which severely degrades the performance of current search systems. To improve the descriptive powers of social tags, a fundamental issue is tag relevance learning, which concerns how to interpret the relevance of a tag with respect to the contents of an image effectively. In this paper, we investigate the problem from a new perspective of learning to rank, and develop a novel approach …


Efficient Motif Discovery In Spatial Trajectories Using Discrete Fréchet Distance, Bo Tang, Man Lung Yiu, Kyriakos Mouratidis, Kai Wang Mar 2017

Efficient Motif Discovery In Spatial Trajectories Using Discrete Fréchet Distance, Bo Tang, Man Lung Yiu, Kyriakos Mouratidis, Kai Wang

Research Collection School Of Computing and Information Systems

The discrete Fréchet distance (DFD) captures perceptual and geographical similarity between discrete trajectories. It has been successfully adopted in a multitude of applications, such as signature and handwriting recognition, computer graphics, as well as geographic applications. Spatial applications, e.g., sports analysis, traffic analysis, etc. require discovering the pair of most similar subtrajectories, be them parts of the same or of different input trajectories.The identified pair of subtrajectories is called a motif.The adoption of DFD as the similarity measure in motif discovery,although semantically ideal, is hindered by the high computational complexity of DFD calculation. In this paper, we propose a suite …


Scalable Image Retrieval By Sparse Product Quantization, Qingqun Ning, Jianke Zhu, Zhiyuan Zhong, Steven C. H. Hoi, Chun Chen Mar 2017

Scalable Image Retrieval By Sparse Product Quantization, Qingqun Ning, Jianke Zhu, Zhiyuan Zhong, Steven C. H. Hoi, Chun Chen

Research Collection School Of Computing and Information Systems

Fast approximate nearest neighbor (ANN) search technique for high-dimensional feature indexing and retrieval is the crux of large-scale image retrieval. A recent promising technique is product quantization, which attempts to index high-dimensional image features by decomposing the feature space into a Cartesian product of low-dimensional subspaces and quantizing each of them separately. Despite the promising results reported, their quantization approach follows the typical hard assignment of traditional quantization methods, which may result in large quantization errors, and thus, inferior search performance. Unlike the existing approaches, in this paper, we propose a novel approach called sparse product quantization (SPQ) to encoding …


Probabilistic Public Key Encryption For Controlled Equijoin In Relational Databases, Yujue Wang, Hwee Hwa Pang Mar 2017

Probabilistic Public Key Encryption For Controlled Equijoin In Relational Databases, Yujue Wang, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

We present a public key encryption scheme for relational databases (PKDE) that allows the owner to control the execution of cross-relation joins on an outsourced server. The scheme allows anyone to deposit encrypted records in a database on the server. Thereafter, the database owner may authorize the server to join any two relations to identify matching records across them, while preventing self-joins that would reveal information on records that are unmatched in the join. The security of our construction is formally proved in the random oracle model based on the computational bilinear Diffie-Hellman assumption. Specifically, before a relation is joined, …


Dark Hazard: Large-Scale Discovery Of Unknown Hidden Sensitive Operations In Android Apps, Xiaorui Pan, Xueqiang Wang, Yue Duan, Xiaofeng Wang, Heng Yin Mar 2017

Dark Hazard: Large-Scale Discovery Of Unknown Hidden Sensitive Operations In Android Apps, Xiaorui Pan, Xueqiang Wang, Yue Duan, Xiaofeng Wang, Heng Yin

Research Collection School Of Computing and Information Systems

Hidden sensitive operations (HSO) such as stealing privacy user data upon receiving an SMS message are increasingly utilized by mobile malware and other potentially-harmful apps (PHAs) to evade detection. Identification of such behaviors is hard, due to the challenge in triggering them during an app’s runtime. Current static approaches rely on the trigger conditions or hidden behaviors known beforehand and therefore cannot capture previously unknown HSO activities. Also these techniques tend to be computationally intensive and therefore less suitable for analyzing a large number of apps. As a result, our understanding of real-world HSO today is still limited, not to …


Intelligent Web Crawler For Semantic Search Engine, Shujia Zhang Feb 2017

Intelligent Web Crawler For Semantic Search Engine, Shujia Zhang

Master's Projects

A Semantic Search Engine (SSE) is a program that produces semantic-oriented concepts from the Internet. A web crawler is the front end of our SSE; its primary goal is to supply important and necessary information to the data analysis component of SSE. The main function of the analysis component is to produce the concepts (moderately frequent finite sequences of keywords) from the input; it uses some variants of TF-IDF as a primary tool to remove stop words. However, it is a very expensive way to filter out stop words using the idea of TF-IDF. The goal of this project is …


Modeling Adoption Dynamics In Social Networks, Minh Duc Luu Feb 2017

Modeling Adoption Dynamics In Social Networks, Minh Duc Luu

Dissertations and Theses Collection

This dissertation studies the modeling of user-item adoption dynamics where an item can be an innovation, a piece of contagious information or a product. By “adoption dynamics” we refer to the process of users making decision choices to adopt items based on a variety of user and item factors. In the context of social networks, “adoption dynamics” is closely related to “item diffusion”. When a user in a social network adopts an item, she may influence her network neighbors to adopt the item. Those neighbors of her who adopt the item then continue to trigger more adoptions. As this progress …


Discovering Burst Patterns Of Burst Topic In Twitter, Guozhong Dong, Wu Yang, Feida Zhu, Wei Wang Feb 2017

Discovering Burst Patterns Of Burst Topic In Twitter, Guozhong Dong, Wu Yang, Feida Zhu, Wei Wang

Research Collection School Of Computing and Information Systems

Twitter has become one of largest social networks for users to broadcast burst topics. There have been many studies on how to detect burst topics. However, mining burst patterns in burst topics has not been solved by the existing works. In this paper, we investigate the problem of mining burst patterns of burst topic in Twitter. A burst topic user graph model is proposed, which can represent the topology structure of burst topic propagation across a large number of Twitter users. Based on the model, hierarchical clustering is applied to cluster burst topics and reveal burst patterns from the macro …


Active Video Summarization: Customized Summaries Via On-Line Interaction With The User, Ana Garcia Del Molino, Xavier Boix, Joo-Hwee Lim, Ah-Hwee Tan Feb 2017

Active Video Summarization: Customized Summaries Via On-Line Interaction With The User, Ana Garcia Del Molino, Xavier Boix, Joo-Hwee Lim, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

To facilitate the browsing of long videos, automatic video summarization provides an excerpt that represents its content. In the case of egocentric and consumer videos, due to their personal nature, adapting the summary to specific user’s preferences is desirable. Current approaches to customizable video summarization obtain the user’s preferences prior to the summarization process. As a result, the user needs to manually modify the summary to further meet the preferences. In this paper, we introduce Active Video Summarization (AVS), an interactive approach to gather the user’s preferences while creating the summary. AVS asks questions about the summary to update it …


Harnessing Twitter To Support Serendipitous Learning Of Developers, Abhabhisheksh Sharma, Yuan Tian, Agus Sulistya, David Lo, Aiko Yamashita Feb 2017

Harnessing Twitter To Support Serendipitous Learning Of Developers, Abhabhisheksh Sharma, Yuan Tian, Agus Sulistya, David Lo, Aiko Yamashita

Research Collection School Of Computing and Information Systems

Developers often rely on various online resources, such as blogs, to keep themselves up-to-date with the fast pace at which software technologies are evolving. Singer et al. found that developers tend to use channels such as Twitter to keep themselves updated and support learning, often in an undirected or serendipitous way, coming across things that they may not apply presently, but which should be helpful in supporting their developer activities in future. However, identifying relevant and useful articles among the millions of pieces of information shared on Twitter is a non-trivial task. In this work to support serendipitous discovery of …


Unsupervised Visual Hashing With Semantic Assistant For Content-Based Image Retrieval, Lei Zhu, Jialie Shen, Liang Xie, Zhiyong Cheng Feb 2017

Unsupervised Visual Hashing With Semantic Assistant For Content-Based Image Retrieval, Lei Zhu, Jialie Shen, Liang Xie, Zhiyong Cheng

Research Collection School Of Computing and Information Systems

As an emerging technology to support scalable content-based image retrieval (CBIR), hashing has recently received great attention and became a very active research domain. In this study, we propose a novel unsupervised visual hashing approach called semantic-assisted visual hashing (SAVH). Distinguished from semi-supervised and supervised visual hashing, its core idea is to effectively extract the rich semantics latently embedded in auxiliary texts of images to boost the effectiveness of visual hashing without any explicit semantic labels. To achieve the target, a unified unsupervised framework is developed to learn hash codes by simultaneously preserving visual similarities of images, integrating the semantic …


Maximizing The Probability Of Arriving On Time: A Practical Q-Learning Method, Zhiguang Cao, Hongliang Guo, Jie Zhang, Frans Oliehoek, Ulrich Fastenrath Feb 2017

Maximizing The Probability Of Arriving On Time: A Practical Q-Learning Method, Zhiguang Cao, Hongliang Guo, Jie Zhang, Frans Oliehoek, Ulrich Fastenrath

Research Collection School Of Computing and Information Systems

The stochastic shortest path problem is of crucial importance for the development of sustainable transportation systems. Existing methods based on the probability tail model seek for the path that maximizes the probability of arriving at the destination before a deadline. However, they suffer from low accuracy and/or high computational cost. We design a novel Q-learning method where the converged Q-values have the practical meaning as the actual probabilities of arriving on time so as to improve accuracy. By further adopting dynamic neural networks to learn the value function, our method can scale well to large road networks with arbitrary deadlines. …


Why And How Developers Fork What From Whom In Github, Jing Jiang, David Lo, Jiahuan He, Xin Xia, Pavneet Singh Kochhar, Li Zhang Feb 2017

Why And How Developers Fork What From Whom In Github, Jing Jiang, David Lo, Jiahuan He, Xin Xia, Pavneet Singh Kochhar, Li Zhang

Research Collection School Of Computing and Information Systems

Forking is the creation of a new software repository by copying another repository. Though forking is controversial in traditional open source software (OSS) community, it is encouraged and is a built-in feature in GitHub. Developers freely fork repositories, use codes as their own and make changes. A deep understanding of repository forking can provide important insights for OSS community and GitHub. In this paper, we explore why and how developers fork what from whom in GitHub. We collect a dataset containing 236,344 developers and 1,841,324 forks. We make surveys, and analyze programming languages and owners of forked repositories. Our main …


Attribute-Based Secure Messaging In The Public Cloud, Zhi Yuan Poh, Hui Cui, Robert H. Deng, Yingjiu Li Feb 2017

Attribute-Based Secure Messaging In The Public Cloud, Zhi Yuan Poh, Hui Cui, Robert H. Deng, Yingjiu Li

Research Collection School Of Computing and Information Systems

Messaging systems operating within the public cloud are gaining popularity. To protect message confidentiality from the public cloud including the public messaging servers, we propose to encrypt messages in messaging systems using Attribute-Based Encryption (ABE). ABE is an one-to-many public key encryption system in which data are encrypted with access policies and only users with attributes that satisfy the access policies can decrypt the ciphertexts, and hence is considered as a promising solution for realizing expressive and fine-grained access control of encrypted data in public servers. Our proposed system, called Attribute-Based Secure Messaging System with Outsourced Decryption (ABSM-OD), has three …


Using An Online Tutorial To Teach Rea Data Modeling In Accounting Information Systems Courses, Poh Sun Seow, Pan, Gary Feb 2017

Using An Online Tutorial To Teach Rea Data Modeling In Accounting Information Systems Courses, Poh Sun Seow, Pan, Gary

Research Collection School Of Accountancy

Online learning has been gaining widespread adoption due to its successin enhancing student-learning outcomes and improving student t academicperformance. This paper describes an online tutorial to teach resource-event-agent(REA) data modeling in an undergraduate accounting information systems course.The REA online tutorial reflects a self-study application designed to helpstudents improve their understanding of the REA data model. As such, thetutorial acts as a supplement to lectures by reinforcing the concepts andincorporating practices to assess student understanding. Instructors can accessthe REA online tutorial at http://smu.asg/rea. An independent survey by the University’sCentre for Teaching Excellence found a significant increase in students’perceived knowledge of REA …


Soal: Second-Order Online Active Learning, Shuji Hao, Peilin Zhao, Jing Lu, Steven C. H. Hoi, Chunyan Miao, Chi Zhang Feb 2017

Soal: Second-Order Online Active Learning, Shuji Hao, Peilin Zhao, Jing Lu, Steven C. H. Hoi, Chunyan Miao, Chi Zhang

Research Collection School Of Computing and Information Systems

This paper investigates the problem of online active learning for training classification models from sequentially arriving data. This is more challenging than conventional online learning tasks since the learner not only needs to figure out how to effectively update the classifier but also needs to decide when is the best time to query the label of an incoming instance given limited label budget. The existing online active learning approaches are often based on first-order online learning methods which generally fall short in slow convergence rate and suboptimal exploitation of available information when querying the labeled data. To overcome the limitations, …


Recurrent Neural Networks With Auxiliary Labels For Cross-Domain Opinion Target Extraction, Ying Ding, Jianfei Yu, Jing Jiang Feb 2017

Recurrent Neural Networks With Auxiliary Labels For Cross-Domain Opinion Target Extraction, Ying Ding, Jianfei Yu, Jing Jiang

Research Collection School Of Computing and Information Systems

Opinion target extraction is a fundamental task in opinion mining. In recent years, neural network based supervised learning methods have achieved competitive performance on this task. However, as with any supervised learning method, neural network based methods for this task cannot work well when the training data comes from a different domain than the test data. On the other hand, some rule-based unsupervised methods have shown to be robust when applied to different domains. In this work, we use rule-based unsupervised methods to create auxiliary labels and use neural network models to learn a hidden representation that works well for …


Streaming Classification With Emerging New Class By Class Matrix Sketching, Xin Mu, Feida Zhu, Juan Du, Ee-Peng Lim, Zhi-Hua Zhou Feb 2017

Streaming Classification With Emerging New Class By Class Matrix Sketching, Xin Mu, Feida Zhu, Juan Du, Ee-Peng Lim, Zhi-Hua Zhou

Research Collection School Of Computing and Information Systems

Streaming classification with emerging new class is an important problem of great research challenge and practical value. In many real applications, the task often needs to handle large matrices issues such as textual data in the bag-of-words model and large-scale image analysis. However, the methodologies and approaches adopted by the existing solutions, most of which involve massive distance calculation, have so far fallen short of successfully addressing a real-time requested task. In this paper, the proposed method dynamically maintains two low-dimensional matrix sketches to 1) detect emerging new classes; 2) classify known classes; and 3) update the model in the …


Detecting Similar Repositories On Github, Yun Zhang, David Lo, Pavneet Singh Kochhar, Xin Xia, Quanlai Li, Jianling Sun Feb 2017

Detecting Similar Repositories On Github, Yun Zhang, David Lo, Pavneet Singh Kochhar, Xin Xia, Quanlai Li, Jianling Sun

Research Collection School Of Computing and Information Systems

GitHub contains millions of repositories among which many are similar with one another (i.e., having similar source codes or implementing similar functionalities). Finding similar repositories on GitHub can be helpful for software engineers as it can help them reuse source code, build prototypes, identify alternative implementations, explore related projects, find projects to contribute to, and discover code theft and plagiarism. Previous studies have proposed techniques to detect similar applications by analyzing API usage patterns and software tags. However, these prior studies either only make use of a limited source of information or use information not available for projects on GitHub. …


Collaboration Trumps Homophily In Urban Mobile Crowd-Sourcing, Thivya Kandappu, Archan Misra, Randy Tandriansyah Daratan Feb 2017

Collaboration Trumps Homophily In Urban Mobile Crowd-Sourcing, Thivya Kandappu, Archan Misra, Randy Tandriansyah Daratan

Research Collection School Of Computing and Information Systems

This paper establishes the power of dynamic collaborative task completion among workers for urban mobile crowdsourcing. Collaboration is defined via the notion of peer referrals, whereby a worker who has accepted a location-specific task, but is unlikely to visit that location, offloads the task to a willing friend. Such a collaborative framework might be particularly useful for task bundles, especially for bundles that have higher geographic dispersion. The challenge, however, comes from the high similarity observed in the spatiotemporal pattern of task completion among friends. Using extensive real-world crowd-sourcing studies conducted over 7 weeks and 1000+ workers on a campus-based …


Crowdsensing And Analyzing Micro-Event Tweets For Public Transportation Insights, Thoong Hoang, Pei Hua (Xu Peihua) Cher, Philips Kokoh Prasetyo, Ee-Peng Lim Feb 2017

Crowdsensing And Analyzing Micro-Event Tweets For Public Transportation Insights, Thoong Hoang, Pei Hua (Xu Peihua) Cher, Philips Kokoh Prasetyo, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Efficient and commuter friendly public transportation system is a critical part of a thriving and sustainable city. As cities experience fast growing resident population, their public transportation systems will have to cope with more demands for improvements. In this paper, we propose a crowdsensing and analysis framework to gather and analyze realtime commuter feedback from Twitter. We perform a series of text mining tasks identifying those feedback comments capturing bus related micro-events; extracting relevant entities; and, predicting event and sentiment labels. We conduct a series of experiments involving more than 14K labeled tweets. The experiments show that incorporating domain knowledge …