Open Access. Powered by Scholars. Published by Universities.®

Models and Methods Commons™

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 30 of 35

Full-Text Articles in Models and Methods

Textual Data Science With R By Monica Becue-Bertaut, Kenneth Benoit Oct 2021

Textual Data Science With R By Monica Becue-Bertaut, Kenneth Benoit

Research Collection School of Social Sciences

Textual Data Science With R targets an important and rela-tively understudied area of data science: the statistical analysisof largely unstructured data in the form of natural languagetext. Using examples spanning fields such as free-form sur-vey responses, bibliographies, and speeches, the book presentsmulti-dimensional methods for mining patterns and insightsfrom textual data. Beginning with a practical and conceptualoverview of textual data and how to pre-preprocess and struc-ture this data, the book proceeds to explain the framework ofcorrespondence analysis and its application to textual data. Itthen discusses two other major approaches: clustering and afocus on cluster features, including characteristic words, andmultiple factor analysis. …


Text As Data: An Overview, Kenneth Benoit Dec 2020

Text As Data: An Overview, Kenneth Benoit

Research Collection School of Social Sciences

When it comes to textual data, the fields of political science and international relations face a genuine embarrassment of riches. Never before has so much text been so readily available on such a wide variety of topics that concern our discipline. Legislative debates, party manifestos, committee transcripts, candidate and other political speeches, lobbying documents, court opinions, laws – not only are all recorded and published today, but in many cases this is in a readily available form that is easily converted into structured data for systematic analysis.


“Daughter” As A Positionality And The Gendered Politics Of Taking Parents Into The Field, Menusha De Silva, Kanchan Gandhi Dec 2019

“Daughter” As A Positionality And The Gendered Politics Of Taking Parents Into The Field, Menusha De Silva, Kanchan Gandhi

Research Collection School of Social Sciences

Research on gendered politics of the field has delved into the practices of accompaniment and its implications on research and knowledge production, particularly through the case of researchers’ children and partners. In comparison, the tendency to seek assistance from parents is neglected within the scholarship. Drawing on the PhD fieldwork experiences of two researchers in their “native” country, specifically a Sri Lankan researcher conducting fieldwork in Sri Lanka and a North Indian scholar researching in South India, the paper reveals parents’ contribution to the research process, in terms of enhancing researcher credibility, facilitating contact‐making and access, and providing emotional and …


Measuring And Explaining Political Sophistication Through Textual Complexity, Kenneth Benoit, Kevin Munger, Arthur Spirling Apr 2019

Measuring And Explaining Political Sophistication Through Textual Complexity, Kenneth Benoit, Kevin Munger, Arthur Spirling

Research Collection School of Social Sciences

Political scientists lack domain-specific measures for the purpose of measuring the sophistication of political communication. We systematically review the shortcomings of existing approaches, before developing a new and better method along with software tools to apply it. We use crowdsourcing to perform thousands of pairwise comparisons of text snippets and incorporate these results into a statistical model of sophistication. This includes previously excluded features such as parts of speech and a measure of word rarity derived from dynamic term frequencies in the Google Books data set. Our technique not only shows which features are appropriate to the political domain and …


Process-Tracing Research Designs: A Practical Guide, Jacob Ricks, Amy H. Liu Oct 2018

Process-Tracing Research Designs: A Practical Guide, Jacob Ricks, Amy H. Liu

Research Collection School of Social Sciences

Process-tracing has grown in popularity among qualitative researchers. However, unlike statistical models and estimators—or even other topics in qualitative methods—process-tracing is largely bereft of guidelines, especially when it comes to teaching. We address this shortcoming by providing a step-by-step checklist for developing a research design to use process-tracing as a valid and substantial tool for hypothesis testing. This practical guide should be of interest for both research application and instructional purposes. An online appendix containing multiple examples facilitates teaching of the method.


Conceptualizing And Measuring Women’S Political Leadership: From Presence To Balance, Devin K. Joshi, Ryan Goehrung Sep 2018

Conceptualizing And Measuring Women’S Political Leadership: From Presence To Balance, Devin K. Joshi, Ryan Goehrung

Research Collection School of Social Sciences

This article conceptualizes an innovative understanding and measurement of women's political leadership, theoretically justifies its application, and analyzes contemporary variation in its patterns through comparative case studies. In recent years, scholars of comparative government have studied with great interest the election of female prime ministers and presidents (e.g., Derichs and Thompson 2013; Jalalzai 2013) and cross-national variation in female members of parliaments (MPs) and cabinets (e.g., Bauer and Tremblay 2011; Paxton and Hughes 2017; Suraj, Scherpereel, and Adams 2014). Yet, when it comes to regions beyond Europe and the Americas, comparative empirical analysis of women's political leadership (WPL) across national-level …


Conceptualizing And Measuring Women’S Political Leadership: From Presence To Balance, Devin K. Joshi, Ryan Goehrung Sep 2018

Conceptualizing And Measuring Women’S Political Leadership: From Presence To Balance, Devin K. Joshi, Ryan Goehrung

Research Collection School of Social Sciences

This article conceptualizes an innovative understanding and measurement of women's political leadership, theoretically justifies its application, and analyzes contemporary variation in its patterns through comparative case studies. In recent years, scholars of comparative government have studied with great interest the election of female prime ministers and presidents (e.g., Derichs and Thompson 2013; Jalalzai 2013) and cross-national variation in female members of parliaments (MPs) and cabinets (e.g., Bauer and Tremblay 2011; Paxton and Hughes 2017; Suraj, Scherpereel, and Adams 2014). Yet, when it comes to regions beyond Europe and the Americas, comparative empirical analysis of women's political leadership (WPL) across national-level …


Text Analysis In R, Kasper Welbers, Wouter Van Atteveldt, Kenneth Benoit Nov 2017

Text Analysis In R, Kasper Welbers, Wouter Van Atteveldt, Kenneth Benoit

Research Collection School of Social Sciences

Computational text analysis has become an exciting research field with many applications in communication research. It can be a difficult method to apply, however, because it requires knowledge of various techniques, and the software required to perform most of these techniques is not readily available in common statistical software packages. In this teacher’s corner, we address these barriers by providing an overview of general steps and operations in a computational text analysis project, and demonstrate how each step can be performed using the R statistical software. As a popular open-source platform, R has an extensive user community that develops and …


Predicting The Brexit Vote By Tracking And Classifying Public Opinion Using Twitter Data, Julio C. Amador Diaz Lopez, Sofia Collignon-Delmar, Kenneth Benoit, Akitaka Matsuo Oct 2017

Predicting The Brexit Vote By Tracking And Classifying Public Opinion Using Twitter Data, Julio C. Amador Diaz Lopez, Sofia Collignon-Delmar, Kenneth Benoit, Akitaka Matsuo

Research Collection School of Social Sciences

We use 23M Tweets related to the EU referendum in the UK to predict the Brexit vote. In particular, we use user-generated labels known as hashtags to build training sets related to the Leave/Remain campaign. Next, we train SVMs in order to classify Tweets. Finally, we compare our results to Internet and telephone polls. This approach not only allows to reduce the time of hand-coding data to create a training set, but also achieves high level of correlations with Internet polls. Our results suggest that Twitter data may be a suitable substitute for Internet polls and may be a useful …


Street-Level Bureaucrats And Irrigation Policy Reform In Southeast Asia, Jacob I. Ricks Apr 2017

Street-Level Bureaucrats And Irrigation Policy Reform In Southeast Asia, Jacob I. Ricks

Research Collection School of Social Sciences

Policy reforms are difficult for developing states, especially when they are meant to improve cooperation and collaboration between private citizens and state officials, such as in the case of education, health care provision, business-state relations, and policing. A large part of this challenge is that the policy reforms required for coproduction of services necessitate development of state capacity in new directions. Using the substantive issue of irrigation reforms, especially those aimed at improving service provision and farmer participation, I identify three lessons for reformers regarding the implementation of policy for the coproduction of services. Drawing on extensive fieldwork in Thailand …


Estimating Intra-Party Preferences: Comparing Speeches To Votes, Daniel Schwarz, Denise Traber, Kenneth Benoit Apr 2017

Estimating Intra-Party Preferences: Comparing Speeches To Votes, Daniel Schwarz, Denise Traber, Kenneth Benoit

Research Collection School of Social Sciences

Well-established methods exist for measuring party positions, but reliable means for estimating intra-party preferences remain underdeveloped. While most efforts focus on estimating the ideal points of individual legislators based on inductive scaling of roll call votes, this data suffers from two problems: selection bias due to unrecorded votes and strong party discipline, which tends to make voting a strategic rather than a sincere indication of preferences. By contrast, legislative speeches are relatively unconstrained, as party leaders are less likely to punish MPs for speaking freely as long as they vote with the party line. Yet, the differences between roll call …


Crowd-Sourced Text Analysis: Reproducible And Agile Production Of Political Data, Kenneth Benoit, Drew Conway, Benjamin E. Lauderdale, Michael Laver, Slava Mikhaylov May 2016

Crowd-Sourced Text Analysis: Reproducible And Agile Production Of Political Data, Kenneth Benoit, Drew Conway, Benjamin E. Lauderdale, Michael Laver, Slava Mikhaylov

Research Collection School of Social Sciences

Empirical social science often relies on data that are not observed in the field, but are transformed into quantitative variables by expert researchers who analyze and interpret qualitative raw sources. While generally considered the most valid way to produce data, this expert-driven process is inherently difficult to replicate or to assess on grounds of reliability. Using crowd-sourcing to distribute text for reading and interpretation by massive numbers of nonexperts, we generate results comparable to those using experts to read and interpret the same texts, but do so far more quickly and flexibly. Crucially, the data we collect can be reproduced …


Debate On Bernard Yack’S Book Nationalism And The Moral Psychology Of Community, Jonathan Hearn, Chandran Kukathas, David Miller, Bernard Yack Jun 2014

Debate On Bernard Yack’S Book Nationalism And The Moral Psychology Of Community, Jonathan Hearn, Chandran Kukathas, David Miller, Bernard Yack

Research Collection School of Social Sciences

Bernard Yack's Nationalism and the Moral Psychology of Community (2012) was the subject of the eighth in a long‐running series of debates hosted by the Association for the Study of Ethnicity and Nationalism (ASEN) and the journal Nations and Nationalism in November of 2013. These debates bring together the authors of recent important works in the study of nationalism and ethnicity with appropriate scholars to explore the questions they have provoked. On this occasion Professor Yack was joined by Professor Chandran Kukathas and Professor David Miller.


Validating Estimates Of Latent Traits From Textual Data Using Human Judgment As A Benchmark, Will Lowe, Kenneth Benoit Jun 2013

Validating Estimates Of Latent Traits From Textual Data Using Human Judgment As A Benchmark, Will Lowe, Kenneth Benoit

Research Collection School of Social Sciences

Automated and statistical methods for estimating latent political traits and classes from textual data hold great promise, because virtually every political act involves the production of text. Statistical models of natural language features, however, are heavily laden with unrealistic assumptions about the process that generates these data, including the stochastic process of text generation, the functional link between political variables and observed text, and the nature of the variables (and dimensions) on which observed text should be conditioned. While acknowledging statistical models of latent traits to be "wrong," political scientists nonetheless treat their results as sufficiently valid to be useful. …


Coder Reliability And Misclassification In The Human Coding Of Party Manifestos, Slava Mikhaylov, Michael Laver, Kenneth Benoit Dec 2012

Coder Reliability And Misclassification In The Human Coding Of Party Manifestos, Slava Mikhaylov, Michael Laver, Kenneth Benoit

Research Collection School of Social Sciences

The Comparative Manifesto Project (CMP) provides the only time series of estimated party policy positions in political science and has been extensively used in a wide variety of applications. Recent work (e.g., Benoit, Laver, and Mikhaylov 2009; Klingemann et al. 2006) focuses on nonsystematic sources of error in these estimates that arise from the text generation process. Our concern here, by contrast, is with error that arises during the text coding process since nearly all manifestos are coded only once by a single coder. First, we discuss reliability and misclassification in the context of hand-coded content analysis methods. Second, we …


Natural Sentences As Valid Units For Coded Political Texts, Thomas Daubler, Kenneth Benoit, Slava Mikhaylov, Michael Laver Oct 2012

Natural Sentences As Valid Units For Coded Political Texts, Thomas Daubler, Kenneth Benoit, Slava Mikhaylov, Michael Laver

Research Collection School of Social Sciences

A rapidly growing area in political science has focused on perfecting techniques to treat politicaltext as ‘data’, usually for the purposes of estimating latent traits such as left–right political policypositions.1 More traditional approaches have applied classical content analysis to categorize sub-unitsof political text, such as sentences in manifestos. Prominent examples of this latter approach includethe thirty-year old Comparative Manifestos Project and the Policy Agendas Project.2 ‘Text as data’approaches use machines to convert text to quantitative information and use statistical tools to makeinferences about characteristics of the author of the text. Content analysis schemes use humans to readtextual sub-units and assign …


How To Scale Coded Text Units Without Bias: A Response To Gemenis, Kenneth Benoit, Michael Laver, Will Lowe, Slava Mikhaylov Sep 2012

How To Scale Coded Text Units Without Bias: A Response To Gemenis, Kenneth Benoit, Michael Laver, Will Lowe, Slava Mikhaylov

Research Collection School of Social Sciences

Coding non-manifesto documents as if they were genuine policy platforms produced at election time clearly raises serious issues with error when these codings are used in the standard manner to estimate left-right policy positions. In addition to the long term solution of improving the document base of the Manifesto Project identified by Gemenis (2012), we argue that immediate gains in manifesto-based estimates of policy positions can be realised by using the confrontational logit scales from Lowe et al. (2011), which addresses the problems of scale content and scale construction that are exacerbated by but not unique to the problems found …


The Dimensionality Of Political Space: Epistemological And Methodological Considerations, Kenneth Benoit, Michael Laver Jun 2012

The Dimensionality Of Political Space: Epistemological And Methodological Considerations, Kenneth Benoit, Michael Laver

Research Collection School of Social Sciences

Spatial characterizations of agents' preferences lie at the heart of many theories of political competition. These give rise to explicitly dimensional interpretations. Parties define and differentiate themselves in terms of substantive policy issues, and the configuration of such issues that is required for a good description of political competition affects how we think substantively about the underlying political space in which parties compete. For this reason a great deal of activity in political science consists of estimating such configurations in particular real settings. We focus on three main issues in this article. First, we discuss the nature of political differences …


Scaling Policy Preferences From Coded Political Texts, Will Lowe, Kenneth Benoit, Slava Mikhaylov, Michael Laver Feb 2011

Scaling Policy Preferences From Coded Political Texts, Will Lowe, Kenneth Benoit, Slava Mikhaylov, Michael Laver

Research Collection School of Social Sciences

Scholars estimating policy positions from political texts typically code words or sentences and then build left-right policy scales based on the relative frequencies of text units coded into different categories. Here we reexamine such scales and propose a theoretically and linguistically superior alternative based on the logarithm of odds-ratios. We contrast this scale with the current approach of the Comparative Manifesto Project (CMP), showing that our proposed logit scale avoids widely acknowledged flaws in previous approaches. We validate the new scale using independent expert surveys. Using existing CMP data, we show how to estimate more distinct policy dimensions, for more …


Challenges For Estimating Policy Preferences: Announcing An Open Access Archive Of Political Documents, Kenneth Benoit, Thomas Brauninger, Marc Debus Sep 2009

Challenges For Estimating Policy Preferences: Announcing An Open Access Archive Of Political Documents, Kenneth Benoit, Thomas Brauninger, Marc Debus

Research Collection School of Social Sciences

We provide a comparative perspective on the contributions of the special issue with regard to their applied methods and findings. In addition, we discuss problems that arise when using ‘wrong’ or at least ‘incorrect’ versions of election manifestos by presenting replications of estimated policy positions of German parties. We show that the latter can result in biased estimates that may affect the outcome of theoretical models. On the basis of those findings, we present the idea of the open access archive polidoc.net to build up a common database for political texts.


Treating Words As Data With Error: Uncertainty In Text Statements Of Policy Positions, Kenneth Benoit, Michael Laver, Slava Mikhaylov Apr 2009

Treating Words As Data With Error: Uncertainty In Text Statements Of Policy Positions, Kenneth Benoit, Michael Laver, Slava Mikhaylov

Research Collection School of Social Sciences

Political text offers extraordinary potential as a source of information about the policy positions of political actors. Despite recent advances in computational text analysis, human interpretative coding of text remains an important source of text-based data, ultimately required to validate more automatic techniques. The profession's main source of cross-national, time-series data on party policy positions comes from the human interpretative coding of party manifestos by the Comparative Manifesto Project (CMP). Despite widespread use of these data, the uncertainty associated with each point estimate has never been available, undermining the value of the dataset as a scientific resource. We propose a …


Benchmarks For Text Analysis: A Response To Budge And Pennings, Kenneth Benoit, Michael Laver Mar 2007

Benchmarks For Text Analysis: A Response To Budge And Pennings, Kenneth Benoit, Michael Laver

Research Collection School of Social Sciences

Budge and Pennings (2007) criticize the “Wordscores” method for computerized content analysis on essentially two grounds. The first is that the best test of Wordscores accuracy is whether it can “reproduce the rich time series produced by the MRG/CMP covering a 50 year period” (Budge and Pennings, 2007: 5), which Budge and Pennings claim it does not do. The second is that Wordscores time series estimates, as implemented by Budge and Pennings, yield very little variation around mean scores for the entire time series. In this brief response we make three simple points.


Estimating Party Policy Positions: Comparing Expert Surveys And Hand-Coded Content Analysis, Kenneth Benoit, Michael Laver Mar 2007

Estimating Party Policy Positions: Comparing Expert Surveys And Hand-Coded Content Analysis, Kenneth Benoit, Michael Laver

Research Collection School of Social Sciences

In this paper we compare estimates of the left-right positions of political parties derived from an expert survey recently completed by the authors with those derived by the Comparative Manifestos Project (CMP) from the content analysis of party manifestos. Having briefly described the expert survey, we first explore the substantive policy content of left and right in the expert survey estimates. We then compare the expert survey to the CMP method on methodological grounds. Third, we directly compare the expert survey results to the CMP results for the most recent time period available, revealing some agreement but also numerous inconsistencies …


Who? Whom? Reparations And The Problem Of Agency, Chandran Kukathas Sep 2006

Who? Whom? Reparations And The Problem Of Agency, Chandran Kukathas

Research Collection School of Social Sciences

If a person is wronged, whether by a physical violation of his person or by having his property unjustly taken, or even by the besmirching of his reputation, he is, most people agree, entitled to some form of compensation or restitution from the person or persons responsible for the wrong. What form the reparation should take, and how great it should be, are sometimes difficult problems, but this does not change the fact that something is owed and someone must be held to account. If a restaurant goes bust because a supplier fails to fulfill his commitments and a newspaper …


Multiparty Split-Ticket Voting Estimation As An Ecological Inference Problem, Kenneth Benoit, Michael Laver, Daniela Giannetti Jan 2004

Multiparty Split-Ticket Voting Estimation As An Ecological Inference Problem, Kenneth Benoit, Michael Laver, Daniela Giannetti

Research Collection School of Social Sciences

The estimation of vote splitting in mixed-member electoral systems is a common problem in electoral studies, where the goal of researchers is to estimate individual voter transitions between parties on two different ballots cast simultaneously. Because the ballots are cast separately and secretly, however, voter choice on the two ballots must be recreated from separately tabulated aggregate data. The problem is therefore of one of making ecological inferences. Because of the multiparty contexts normally found where mixed-member electoral rules are used, furthermore, the problem involves large-table (R × C) ecological inference. In this chapter we show how vote-splitting problems in …


Extracting Policy Positions From Political Texts Using Words As Data, Michael Laver, Kenneth Benoit, John Garry May 2003

Extracting Policy Positions From Political Texts Using Words As Data, Michael Laver, Kenneth Benoit, John Garry

Research Collection School of Social Sciences

We present a new way of extracting policy positions from political texts that treats texts not as discourses to be understood and interpreted but rather, as data in the form of words. We compare this approach to previous methods of text analysis and use it to replicate published estimates of the policy positions of political parties in Britain and Ireland, on both economic and social policy dimensions. We “export” the method to a non-English-language environment, analyzing the policy positions of German parties, including the PDS as it entered the former West German party system. Finally, we extend its application beyond …


Estimating Irish Party Policy Positions Using Computer Wordscoring: The 2002 Election: A Research Note, Kenneth Benoit, Michael Laver Jan 2003

Estimating Irish Party Policy Positions Using Computer Wordscoring: The 2002 Election: A Research Note, Kenneth Benoit, Michael Laver

Research Collection School of Social Sciences

Developments in the computerised analysis of political texts now make itmuch easier than before to investigate large volumes of political text inorder to estimate the policy positions of the authors. Previous contentanalyses of party manifestos, for example, have relied on the hand codingof texts (Budge et al., 1987; Laver and Budge, 1992; Klingeman et al.,1994; Budge et al., 2001), or on dictionary-based computer codingtechniques (Laver and Garry, 2000; Kleinnijenhuis and Pennings, 2001;Garry, 2001; de Vries et al., 2001; Bara, 2001). Such analyses, even thosebased on computer coding dictionaries, require heavy human involvement or intervention, creating both a huge resource cost …


The Endogeneity Problem In Electoral Studies: A Critical Re-Examination Of Duverger's Mechanical Effect, Kenneth Benoit Mar 2002

The Endogeneity Problem In Electoral Studies: A Critical Re-Examination Of Duverger's Mechanical Effect, Kenneth Benoit

Research Collection School of Social Sciences

Studies of electoral law consequences typically treat electoral laws as exogenous factors affecting political party systems, even while acknowledging that political parties often tailor electoral institutions to suit their own distributional needs. This study represents a departure from that approach, directly examining one aspect of the endogeneity of electoral systems: the endogeneity of Duverger's ‘mechanical’ effect. Theory clearly posits that the Duvergerian ‘psychological’ effect of electoral rules occurs in anticipation of their reductive mechanical effect, yet in empirical models this endogenous character is typically ignored. In this paper I formalize the two types of Duvergerian effects of electoral laws in …


Locating Tds In Policy Spaces: The Computational Text Analysis Of Dáil Speeches, Michael Laver, Kenneth Benoit Jan 2002

Locating Tds In Policy Spaces: The Computational Text Analysis Of Dáil Speeches, Michael Laver, Kenneth Benoit

Research Collection School of Social Sciences

This article adapts a new technique for the computerised analysis of political texts, previously used to analyse party manifestos, to the analysis of speeches made in a legislature. The benefits of computerised text analysis come from the ability to analyse, for the first time, complex and daunting electronic sources of text, such as the parliamentary record. This allows the systematic estimation of the policy positions of individual political actors, with huge benefits both for theory development and empirical analysis. In this article, the technique is used to analyse all 58 English language speeches made in the October 1991 confidence debate …


Can A Liberal Society Tolerate Illiberal Elements?, Chandran Kukathas Jan 2001

Can A Liberal Society Tolerate Illiberal Elements?, Chandran Kukathas

Research Collection School of Social Sciences

Libertarians believe that all individuals are entitledto live as they choose, free from interference byother persons or by the state. They also believethat in the absence of such interference, whether bygovernment or other agents of the state intent ondesigning or planning for society as a whole, order willnonetheless prevail. Given the freedom to contract andexchange, markets will coordinate the production anddistribution of goods — and indeed do so better thanany other institution can.