scieee AI-readable full text Open interactive document viewer

Construction and analysis of political networks over time via government and me

Garcia-Olano, Diego

Abstract

In this work we present a tool that generates real world political networks from user provided lists of politicians and news sites. We use as input a dataset of current Texas politicians and 6 news sites to illustrate the graphs, tools and maps created by the tool to give users political insight.

Full text

Automated Construction and Analysis of Political Networks via open government and media sources. MASTERS THESIS DIEGO GARCIA-OLANO, Universitat Politècnica de Catalunya Advisors: Marta Arias and Josep Lluís Larriba Pey A joint research project with the DAMA and LARCA groups at Universitat Politècnica de Cataluyna (UPC) – Barcelona Tech Department of Informatics (FIB) Submitted in partial fulfillment of the requirements for the degree of Master in Innovation and Research in Informatics Specialization: Data Mining and Business Intelligence June 30, 2015 ACKNOWLEDGEMENTS I would like to thank Marta Arias for being so gracious with her time and providing me with suggestions, editing, feedback and support through out this process. I would constantly make jokes about how she should be awarded for the number of Masters students she was advising concurrently in addition to myself, but I really wasn’t joking. I would like to thank Josep Lluís “Larri” Larriba Pey for allowing me to work on this project, for inspiring me to work hard on it’s initial phase for a Knight Foundation Media grant and for his support particularly with respect to my mother who was diagnosed with lung cancer near the beginning of this work, and who just yesterday, after months of treatment, was told she is in the clear! I would like to thank everyone from the DAMA and LARCA groups, but particularly Francisco Rodriguez Drumond with whom I brainstormed through out the process and discussed ideas. Additionally, I would like to thank Lluis Belanche for his excellent Machine Learning course, his in-class jokes, love of Prog Rock and for letting me help on his research involving Microbaterial source tracking. I would like to thank my dear friends Flora Lichtman, and Meghan Fergusson for helping make a video used for the grant application that involved a talking cactus wearing a Starburst candy wrapper as a bandana and his female companion, a Texan tumbleweed with boots on who sounds remarkably like my partner Cecilia. I would like to thank Chris Valdez for many reasons; for the initial design work of the Who You Elect tool, and for all that he has done for my mother during the past two years. I would like to thank Lou and Berta of Babelia, my favorite café/bookstore in Barcelona where a good deal of this work took place, for always smiling and being good sports when I tried to explain to them what I was doing. There are many others I would like to thank, particularly some of the other fantastic professors I’ve had the fortune of studying under at the UPC, but I will leave that for another time. This thesis is dedicated to my mother Doctor Lilita Olano for everything, always. ABSTRACT In this work we present a tool that generates real world political networks from user provided lists of politicians and news sites. The tool downloads articles in which the politicians appear amongst the different news sites, processes them, enriches them with data obtained from various open sources and then generates various network visualizations, tools and maps that allow the user to explore and better understand those politicians and their surrounding environments. To demonstrate the capabilities of the tool for use in studying political and media landscapes, we construct a comprehensive list of current Texas politicians, select six news sites that convey a spectrum of representative political viewpoints publishing articles on the state, and examine the results produced by running the system with them as input. We propose a “Combined” co-occurrence distance metric to better assess the strength of the relationship between two actors in a graph and additionally provide automated summarization tools that utilize text-mining techniques to extract topics and issues characterizing the individual politicians. A similar topic modeling technique is also proposed as a novel way of labeling communities that exist within a politician’s “extended” network. Finally we present media centric results of our case study that show who the different news sources publish articles about both from a geographic and individual perspective. 1 TABLE OF CONTENTS ACKNOWLEDGEMENTS …………………………………………………………………………………………………………………….1 ABSTRACT ………………………………………………………………………………………………………………………………………..1 1. INTRODUCTION …………………………………………………………………………………………………………………………….3 1.1. MOTIVATION …………………………………………………………………………………………………………………………….3 1.2. PROBLEM DESCRIPTION AND BACKGROUND ………………………………………………………………………………...3 1.3. OVERVIEW OF SYSTEM: “WHO YOU ELECT” …………………………………………………………………………………..4 1.4. DESCRIPTION OF CASE STUDY: TEXAS POLITICS …………………………………………………………………………….4 2. RELATED WORKS ………………………………………………………………………………………………………………………….5 3. AUTOMATED CONSTRUCTION OF NETWORKS …………………………………………………………………………………7 3.1 ADDING POLITICIANS FROM OPEN GOVERNMENT SOURCES …………………………………………………………….7 3.2 ADDING NEWS SOURCES ……………………………………………………………………………………………………………..8 3.3 SET UP FOR DATA ACQUISITION BY TEMPLATE MODIFICATION ………………………………………………………..8 3.4 RUNNING AND STORING THE WEB SEARCH RESULTS FOR EACH ACTIVE ENTITY ………………………………..9 3.5 PROCESSING ARTICLE RESULTS PER ACTIVE ENTITY …………………………………………………………………….10 3.6 GENERATE INDIVIDUAL STAR AND EXTENDED GRAPHS FOR EACH INDIVIDUAL ………………………………..12 3.7 ON PREPROCESSING THE CANDIDATE LIST OF POLITICIANS …………………………………………………………..12 4. OVERVIEW OF WHO YOU ELECT VISUALIZATION TOOLS ………………………………………………………………..13 4.1 TABLE OF CONTENTS VIEW ………………………………………………………………………………………………………..13 4.2 MAPS OF TEXAS HOUSE, SENATE AND FEDERAL CONGRESSIONAL DISTRICTS …………………………………..13 4.3 TEXAS COMMITTIES VIEW …………………………………………………………………………………………………………..15 4.4 INDIVIDUAL “STAR” NETWORK VIEW …………………………………………………………………………………………...15 4.4.1 CENTRAL ENTITY VIEW ………………………………………………………………………………………………………..16 4.4.2 SIDE BAR ENTITY ARTICLES TEXT VIEW …………………………………………………………………………………17 4.4.3 TOP ASSOCIATED DISTANCE METRICS ……………………………………………………………………………………17 4.4.4 ARTICLE STATISTICS TEMPORAL VIEW …………………………………………………………………………………..18 4.4.5 “COMBINED” METRIC & OTHER DISTANCE METRICS SUMMARY & COMPARISON VIEW …………………..19 4.5 EXTENDED VIEW WITH COMMUNITY DETECTION ……………………………………………………………………………21 4.5.1 ENTITY-ENTITY GRAPH EXPANSION VIEW ………………………………………………………………………………22 4.5.2 COMMUNITY EXPANSION VIEW ……………………………………………………………………………………………..25 4.5.3 ADDITONAL CENTRALITY MEASURE TOOLS AND VISUALIZATION OPTIONS …………………………………25 4.5.4 COMMUNITY ANALYSIS VIEW ………………………………………………………………………………………………..29 4.6 MEDIA ANALYSIS TOOLS ……………………………………………………………………………………………………………...31 4.6.1 MEDIA ANALYSIS TABLE VIEWS …………………………………………………………………………………………….31 4.6.2 MEDIA ANALYSIS HEAT MAPS OF TEXAS HOUSE, SENATE AND FEDERAL DISTRICTS …………………….33 5. ADDITIONAL ANALYSES ………………………………………………………………………………………………………………..36 5.1 AUTOMATED SUMMARIZATION OF POLITICIANS ……………………………………………………………………………..36 5.2 AUTOMATED SUMMARIZATION OF COMMUNITIES …………………………………………………………………………..38 5.3 MEDIA CENTRIC RESULTS OF CASE STUDY ……………………………………………………………………………………39 6. CONCLUSIONS & FUTURE WORK …………………………………………………………………………………………………...41 7. REFERENCES ………………………………………………………………………………………………………………………………..43 8. APPENDICES ……………………………………………………………………………………………………………………..…………44 0 1. INTRODUCTION 1.1 Motivation We live in an age of over-information, where we must constantly filter information, process it and make educated decisions be it at home, at work or when we vote. However in regards to voting, we are inundated by the media with news stories on a national level, which should in theory lead to a better informed populace, but more often than not, we only consume media outlets which reaffirm our political beliefs and thus entrench us in them, creating straw men of differing views and increasing polarization of voters. This situation is difficult, but at least manageable when a presidential election occurs, as we need only sort through the information related to the histories, views, and promises of a small field of candidates to make our decision. The situation is however much worse on the local or state level where the opposite occurs. There are vastly more important decisions to be made, and we do not receive enough information regarding all candidates and election races in addition to the aforementioned issue of media bias, both self imposed through the choices we make as media consumers in addition to the biases of the media sources themselves. This overwhelming flood of information results in voters making uninformed or partially incomplete decisions at best or at worst sees them abstaining completely from the process. For instance, the last election for Governor of Texas in 2014 saw the Republican candidate Greg Abbott beat the Democrat one Wendy Davis by a 20% margin, 2.8 million votes versus 1.8 million votes [13]. However one must keep in mind that Texas has approximately 16.7 million eligible voters of whom 4.8 million voted thus giving the contest a dismal 27.5% of eligible voter participation. That number bumps up only to 32.5% if one uses registered voters, but that is still low. Another way of looking at it is to observe that the most powerful executive political position of Texas, a position that directly affects the daily lives and welfare of 27 million Texans was selected by 16% of its possible eligible voters. 1.2 Problem description and background There exists a vast literature on the use of network analysis to study important social and political phenomena [2,3,4,5,23]. This work was in fact inspired in part by a presentation of one such paper describing a novel use of natural language processing and network analysis techniques to describe the network of drug cartels in Mexico[1]. In it, the authors perform text mining, partially by hand and partially automated, on a single book about the subject “Los Señores del Narco” by Anabel Hernandez and then derive a visual network from the actors and links discovered therein as seen in Figure 1.1. Fig. 1.1 Drug Cartels Network from “A literature-based approach to a narco-network”[1] The larger nodes in the figure are the heads of cartels while the links connecting nodes imply a relationship between them; the thicker the link, the more important the relationship. In network analysis, nodes are sometimes referred to as vertices, links are sometimes referred to edges, and 1 3 networks are referred to as graphs. They are used interchangeably and mean the exact same thing. The network derived from the book found 1037 nodes and 6405 links between them with Figure 1.1 being a filtered view highlighting the most important ones. The drug cartel network that was derived from the book was impressive for its visual synthesis of the content in the book. It in no way replaces the book by any stretch of the imagination, but it does provide a simple and concise highlevel view of the actors and entities involved which is highly useful to someone new, but interested in the area. This general idea of combining text mining and network science to both summarize content while allowing for the discovery of interesting relationships is a solid foundation on which to build upon that can handle the enormous, but under utilized, amount of information available in online news. Such processing of largely unstructured text combined with simple, flexible and powerful tools to explore and understand it in a targeted or broad way would be of great use for voters, journalists and researchers alike. The end goal of the work begun with this thesis is to provide a tool that can be both a hosted end solution for users to educate themselves and a generalized framework that may be built upon and deployed by interested parties for entirely new purposes of study. A solution that does not require a data center or anything else that may preclude adoption for financial or code proprietary reasons is also stressed. 1.3 Overview of System In an effort to provide more insight into primarily local and state wide contests, but also including federal elections pertaining to a specific state, we decided to build a system “Who You Elect”1 that could take one or many candidate names as input, along with a list of online new sources, and then retrieve all the articles pertaining to the candidates from them. Conversely the input could be a list of politicians or any group of people a researcher wishes to analyze. The system then using natural language processing, information retrieval and network analysis techniques automatically generates the network of all the politicians, organizations, businesses and locations associated with each candidate inputted. The system makes it easy to add new sources to pull content from and additionally provides two types of visualizations: • An individual close up “star” view that allows a user to view the entities (politicians, businesses, etc.) most associated with a candidate along with the articles and textual context in which they co-occurred. • An “extended” world view that is more of a global view of the network formed from articles pertaining to a candidate which in addition to showing links between himself and associated entities, also shows the links between those entities themselves and thus allows for the detection of communities and other traditional network analysis measures. 1.4 Description of Case Study: Texas Politics For illustrative purposes, we decided to test our system on the political landscape of Texas in 2015. On a state level, the Texas Congress is composed of the Texas House of Representatives and Texas Senate. As of 2015, the state’s congress, or legislature by which it is also referred, is composed of 181 members, 31 senators and 150 representatives respectively. On the federal level within the United States Congress, Texas has 38 total members including 2 senators and 36 representatives. Additionally Texas has 27 other State Level Elected Officials including Governor, Lieutenant Governor, Attorney General, Comptroller of Public Accounts, Commissioner of General Land Office, Commissioner of Agriculture, 3 Railroad Commissioners, 9 Texas Supreme Court Justices, and 9 Court of Criminal Appeals Judges. Thus, in total we have 246 Texas politicians composed of 74 Democrats and 172 Republicans that we will be observing. 1 Who You Elect tool using Texas specific case data at http://www.whoyouelect.com/texas 4 The organization of this document is as follows. Chapter 2 will cover related works and additional background. We will begin by describing the types of graphs the system constructs, along with rational for decisions made, and then look at other works that have focused on deriving networks from unstructured texts. Chapter 3 will describe the back end of the system, namely the methods by which we automatically gather, store and process the articles for the politicians we are studying. Chapter 4 focuses on presenting and explaining the frontend web tools developed and currently available at the WhoYouElect.com; particular attention is paid to demonstrating how interesting relationships may be discovered and verified by using the different tools available. Chapter 5 provides various media-centric statistics on the articles returned from the different new sources used in the case study and provides tools and methods for summarizing content found for the politicians studied via text analysis. Finally Chapter 6 contains a summary of conclusions and future work. 2. RELATED WORKS The types of graphs we will be constructing are undirected heterogeneous networks with weighted edges in political contexts. Heterogeneous in this context just refers to the existence of different node types that will compose the graph. For instance, nodes in our system will be labeled as people, organizations, politicians, locations, bills, or miscellaneous. For an introduction and overview of the state of the art of heterogeneous networks and mining techniques see the following [6,7]. The graphs could be considered to be Social Networks involving political and nonpolitical actors or “noisy” Political Networks due to the inclusion of nodes and relationships not involving politicians, though in the end the distinction is largely semantic and unimportant. Additionally, for the moment we are only considering one relationship type, i.e. at most one edge between each pair of nodes, based on various distance metrics and as such do not find ourselves in the multiplex context which is more adapt for studying complex networks. For an extensive examination of the field, uses, and visualization tools see [15,24]. Adapting the system to include more than one edge type is largely based on being able to create multiple edges between pairs of nodes, by perhaps leveraging linked datasets, and then characterize those relationship types effectively. The former task is largely an information retrieval one usually involving enriching datasets via a knowledge base such as Wikipedia, but usually suffers for queries lacking their own Wiki entries or incomplete data. Two recent works in this area seem promising; one exploring boosting specifically in the case of missing or incomplete linked data involves mining a knowledge base using additional textual context, i.e. “evidences”, for named entity disambiguation [20] and the other constructs a probabilistic model that captures the “popularity” of entities and the distribution of multi-type objects appearing in the textual context of an entity using meta-path constrained random walks over networks to learn weights associated with entity linking [21]. The later task of characterizing relationships between nodes is usually handled by deriving a topic model from the corpus of text available, all of the news articles gathered for a given politician in our case, and assigning the most probable “topic” to an edge based on learned Bayesian probability models [16,17] and the text contexts of the two nodes. A good overview of the idea and the use of Latent Dirichlet Allocation (LDA) or alternatively Latent Semantic Indexing (LSI) to produce topic models is found in [11]. A topic in this context simply refers to a collection or distribution of words that characterizes a predefined or latent concept. For instance, an example topic “Actor” could be defined by the set of words “Hollywood, movie, film, theatre, etc.“ Chang’s paper [16] is particularly interesting as it both learns topics without supervision and assigns them to edges, by constructing “entity topics” and “pair topics” and then considering the two most probably “entity topics” assigned to a pair of nodes along with their most probably “pair topic” in order to final construct a description of their relationship, itself another set of disjoint words. Other approaches leverage the use of “phrases” [17] as opposed to single words to help improve performance, but the basis is the same. The topic model based approach has the advantages that it reduces the dimensionality of the search space since you are now working 2 5 with topic objects as opposed to articles, and the number of topics is intended to be much less than the number of documents, and additionally, it allows for the formation of relationships between entities even if they don’t co-occur in the same document. The disadvantages are that labeling and verifying the resulting topics in itself is a largely manual process [16], the topics themselves can be very noisy and difficult to interpret, and precisely that it moves away from the article centric approach and in doing so creates groupings, similarly to community detection and clustering methods, which themselves impose a structure which may be warranted or not. Because the focus of this initial phase was to construct a working proof of concept and the inclusion of edge labeling via topic modeling would be a nice, but unnecessary step, it has been left for future work. We do some separate textual analysis based on topic modeling however to summarize politicians using their texts, but use it as a ground truth and complementary analysis tool. The use of topic modeling to find potentially interesting relations especially through the use of Congressional bills and their debates as separate topics could be particularly useful. The following recent work describes this idea in the context of Catalan political networks [26]. Similarly the role discovery technique proposed in [19] which is a sort of complement to community detection and uses the structural behavior of nodes to assign them “roles” to aid in the identification of nodes of interest was left for future work. There exist a good deal of prior work that has focused on deriving information networks from unstructured news and other web texts [8,9,10,11]. Two introductions and overviews on the process of deriving networks from text may be found [22,25]. Of the prior works cited some rely on hand crafted networks “extracted through a time and effort consuming manual process based on interviews and questionnaires collected”[9], while others rely on organizations such as the European Media Monitor [8] providing them with access to an article lookup system that while impressive in the breadth of sources available, constrains them to only sources from that list. Additionally and specifically in reference to the European Media Monitor and another considered media aggregation service MIT’s Mediacloud, in the cases where we noticed an overlap in available sources between their listings and the ones we are considering for our case study (the Houston Chronicle for instance), the search results from the original news source internal search engine always returned more results for specific entity queries than either the EMM or Mediacloud service which points to an additional quality assurance weakness1. Other works avoid the actual aspect of retrieving content from news sites by using a point in time snapshot of curated news corpora released through the Linguistic Data Consortium or the New York Times[10, 11]. Similarly, [9] leverages paid-for search engine results, only grabbing the first twenty results for each query and then only utilizing the snippet of text present in Yahoo’s search results page as opposed to all the content within the actual article itself. In the end we were unable to find any work that leveraged the publically available search engines present in most news websites. By utilizing this mechanism and allowing the flexibility to pull content from any news site which fulfills that requirement, our system allows site administrators to create a context in which to search and in doing so curate the content and thus satisfy the needs of different users. Information retrieval aspects of the texts used in the prior works aside, the prior works all use NLP to extract entities and then leverage either similarity metrics based on some combination of entity cooccurrence, textual contexts of and shared between entities, and the correlation of entities and hyper links found in documents [8,9] or topical modeling [10,11] to infer relationships of interest. The textual context presented in [9] is limited in that it does not consider entire documents, rather only text snippets from search engine results, but is robust in its evaluation of metrics quantifying the use of co-occurrences metrics for labeling relationships as positive or negative. Based on the number of results, they produce four metrics of similarity: the Jaccard Coefficient, the Dice Coefficient, Mutual Information (MI) and the Google-based semantic relatedness [29 and evaluated them against a small hand-crafted Policy Network. Although interesting this approach does not scale as an evaluation approach. The works of [27,28] like [8] uses the EMM for content, but are novel in that as opposed to using similarity distance metrics or topic modeling, they also use natural language processing with 1 see Appendix A0 for a sample comparison of a query comparing EMM vs HC result sets. 6 an initial seed of hand crafted syntactic templates to learn syntactic patterns which paraphrase certain predefined relations from articles and uses them to label relationships between entities. The work in [11] is of particular interest and leverages the CATHY framework (Construct A Topical HierarchY), which is a recursive clustering and ranking approach for topical hierarchy generation similar to the idea presented in [16] except with the hierarchical component which allows for relating of topics to their parents “more general” or children “more specific”. As mentioned before however, topic modelling approaches for relationship labeling and as a complementary view for better inclusion of bills in Congress are left for future work. 3. AUTOMATED CONSTRUCTION OF GRAPHS In our case study, we constructed the graphs for 247 active Texan politicians. This political environment was chosen to illustrate the use of the system, but could just as easily have been a list of politicians for any city, state, or country. The system components which deal with the construction of graphs use only Python, some associated open libraries, MongoDB, and occasional Bash scripts, and as such is very light from a technical requirements view point. The general outline of the process for constructing graphs is shown in Figure 3.1 Fig 3.1 General Overview of Graph Construction Process 3.1 Adding Politicians From Open Government Sources We first needed to generate the Texas specific list of congressmen. We leveraged the API for the Sunlight Foundation’s website Openstates.org to obtain a list of both active and inactive members of Texas congress returned as JSON1 and saved them locally. Appendix A contains a sample record the OpenStates API returns for a Texas representative and also a sample record for a Federal Representative returned from the GovTrack.us API. The inactive members API returns a set of 112 prior State representatives and though they will not be analyzed and no articles specific to them will 1 http://openstates.org/api/v1/legislators/?state=tx&active=true http://openstates.org/api/v1/legislators/?state=tx&active=false 3 7 be retrieved, we still will include them as possible domain knowledge to leverage during disambiguation of entity types during the processing of articles stage. Throughout this work “entities” is an umbrella term encompassing any actors of interest including politicians, organizations, bills, etc. Next, we needed the list of state wide elected officials who are not representatives (Governor, Lieutenant Governor, etc.) and made use of the Secretary of State of Texas website1 and script2 to generate a JSON file we then saved locally. Finally, in order to get federal representatives and senators for the state of Texas, we leveraged the API at www.govtrack.us3 and saved the JSON file locally. At this point we have four files referring to the active and inactive state representatives, state level elected officials, and federal level elected officials. We load them all and standardize formatting of fields and save them into our "entities" mongo database via a script4 that also creates a CSV file of only active entities5. 3.2 Adding News Sources: Now that the entities have been added to our database, we select all the news sources to use for our case study. A subset of Texas newspapers with the highest circulation, the Dallas Morning News6, Houston Chronicle7, and the Austin American Statesman8 representing the spectrum of conservative, centrist and progressive areas within Texas were selected along with two sites, the Texas Observer9 and Texas Tribune10 which focus on Texas politics and issues. In addition, the New York Times11 was selected to provide an outside context. This selection of sources was in the end entirely subjective, but made with the intent to present a reasonable mix of representative media consumed within the state of Texas. 3.3. Set up for Data Acquisition By Template Modification: For each desired source we need to gather articles pertaining to our politicians list. We’ve created two template web scrapers, one based on the Python package BeautifulSoup (BS) and another based on the Web.Selenium Python package which utilizes both Chrome and PhantomJS (PS) drivers. Each template version is a folder containing two files: • 1. a file, based on BS or PS, which calls the internal search URL for a given news source, and then saves a JSON file of the article URLs, titles, and dates retrieved. • 2. a file invoked after the first that then grabs the actual article content for each URL obtained from step one and saves it into a separate JSON file, one per article which contains the title, URL, article text, the news-source itself, time, and an identifier. The reason two types of templates are warranted is based on the peculiarities of web scraping and limitations of each. If a website such as the Houston Chronicle practices good web standards and renders site content without the explicit use of JavaScript, since doing so breaks accessibility requirements for non JavaScript browsers including braille reader systems for seeing impaired users, we simply use the Beautiful soup version template to construct our scraper. Unfortunately not all websites render content as such so we take advantage of headless browser technologies via 1 http://www.sos.state.tx.us/elections/voter/elected.shtml 2 http://github.com/diegoolano/thesis/whoyouelect/js/htmltable_to_json.js 3 https://www.govtrack.us/api/v2/role?current=true&state=TX 4 http://github.com/diegoolano/thesis/generate_network/load_entities_from_file_into_db.py 5 http://github.com/diegoolano/thesis/generate_network/active-entities-list.csv 6 http://dallasnews.com 2 http://www.chron.com 3 http://www.statesman.com 4 http://www.texasobserver.org 10 http://www.texastribune.org 11 http://www.nytimes.com 8 On a side note the effect of gerrymandering, particularly at the Federal level, is quite apparent when the maps are view simultaneously; notice the lack of blue in the southwest region. Gerrymandering is “the practice of redrawing legislative boundaries so that the resultant political landscape features built-in electoral advantages for a specific constituency.”[32] Districts are generally redrawn every 5 or 10 ten years depending on the district in question, and the political party in control of that legislature largely gets to decide them thus making the process a remarkably important political affair. For background and interesting work on visualizing and resolving gerrymandering, quantifying its efficacy, and even a theory of optimal partisan gerrymandering see [30,31,32,33,34] . 4.3 Texas Committees View The Committee View, seen in Figure 4.5, is based on two APIs available via OpenStates.org, one that gathers all the committee data available for a state1 and another2 that takes a committee id from the first and returns a list of members who are a part of it. Fig. 4.5 Committee Assignments for the Texas Congress A script was written to automate the downloading of all committee assignments3. Senate members are appointed to committees by the Lieutenant Governor, and the Speaker of the House’s duties include the appointment of chairships and committee membership. The current Speaker of the House is Republican Joe Straus, the Speaker Pro Tempore who fills in in case of absence of the Speaker is Republican Dennis Bonnen, and the current Lieutenant Governor is Republican Dan Patrick. These committee assignments, in the same way spatial proximity of representatives does, provide us with certain associations we expect to find. As an aside, although it may seem that some of the contents are out of date since the “updated date” appears old for certain committees in the visualization, it has been verified that these are the current assignments for the 84th Legislature4 . 4.4 Individual “Star” Network View When “Inner Network” is selected from the “Table of Contents” view, the politician’s processed graph is visually displayed with the entities (ie, people, politicians, organizations, locations, etc.) which have most co-occurred with him being placed closer to his central position. By “most co-occurred” we refer to overall counts independent if the co-occurrences were in the same sentence, near sentences or 1 openstates.org/api/v1/committees/?state=tx 2 http://openstates.org/api/v1/committees/TXC000010/ 3 js/committees/committee-ids.sh 4 http://www.house.state.tx.us/_media/pdf/committee.pdf 15 farther. Figure 4.6 shows the graph produced for Democratic State Representative of District 51, Eddie Rodriguez, who we will be using as a running example in this thesis. Fig. 4.6 Individual Star Network Landing Page for Representative Eddie Rodriguez In the right hand side area we see that 362 articles were obtained from the six sources with the bulk coming from the Houston Chronicle (136), Austin American Statesman (133) and Texas Tribune (60). Additionally we see that 6226 entities were discovered along with the exact break down of counts for each entity types shown. We are presented with the option to filter by date range, filter by which news sources to include, show more, less or all (“Show Full”) nodes on the screen. The “ARTICLE URLS” link at the bottom right displays a pop up window of a sortable table of all 362 articles including their URL, date, number of sentences, number of unique entities, and number of relations created from it. This area is particularly interesting for someone interested in studying the media aspect of which sources reported on an politician and when. The functionality for “SHOW STATS” is of particular interest and is explained in detail in Sections 4.4.2 through 4.4.4. When hovering over a circle on the left hand side, the user sees how many co-occurrences that entity shares with Eddie Rodriguez. Additionally, politicians have their circles shaded with the color of their party affiliation, red for Republican and blue for Democrats. The bottom left yellow icon displays a general user manual. 4.4.1 CENTRAL ENTITY VIEW The results of clicking on the center node Eddie Rodriguez, the politician under study, are seen in Figure 4.6.1. In it we can see the entities’ information, name, party, position, and district number, and additionally we see a map of the district represented. Immediately beneath the map is a link that displays the same pop up window of statistical findings as the “SHOW STATS” link explained in 4.4.3. In addition, we see a sortable table of the findings visualized in the graph with the exception that all entities are shown. Entity types are color coordinated and may be filtered by clicking on the appropriate header. 16 Fig. 4.6.1 Individual Star Network Center View for Representative Eddie Rodriguez 4.4.2 SIDE BAR ENTITY ARTICLES TEXT VIEW Figure 4.7 shows the result when the node for “Tesla” is selected. The right hand side shows that 44 co-occurrences where discovered between Eddie Rodriguez and Tesla in 12 different articles, 3 from the Austin American Statesman (AAS), 3 from the Houston Chronicle (HC), 2 from the Dallas Morning News (DMN) and 4 from the Texas Tribune (TXR). Any of the news source headers can be clicked to display only the articles from that source. The articles are listed in reverse chronological order with the article title being a link to the original article that is colored to match its sources color. Fig. 4.7 View of Co-occurrences between Eddie Rodriguez and Tesla in articles processed from the Austin American Statesman 4.4.3 TOP ASSOCIATED DISTANCE METRICS Figure 4.8 shows the results of clicking on “SHOW STATS” from the landing page or from the right hand side view for when the central Eddie Rodriguez node is clicked. The “Top Associated” tabs along the top of the window each show what are the resulting most associated entities by entity type 17 if different distance metrics are used. For instance, figure 4.8 has the Top Associated “same sentences” tab currently selected, hence highlighted in red, and as such we see a ranked list of the entities which have the most “same sentence” co-occurrences with Eddie Rodriguez separated by entity type (Politicians, Organizations, Person, Location, Bills, Misc). Fig. 4.8 “Show Stats” Screen for Texas Representative Eddie Rodriguez We observe that “Dawna Dukes” is the Politician with the most same sentence occurrences with Eddie Rodriguez where on the Landing Page visualization, which uses the third Top Associated metric; “Tom Craddick” is the Politician with the most same article co-occurrences with Eddie Rodriguez. The fourth Top Associated metric “same article only (not near)” refers to entities that occur the most at a distance from the main politician being studied. This listing can be used as a sort of specialized, local term frequency – inverse document frequency (TF-IDF) measure because it allows the user to observe which entities occur the most at a distance from the politician of study, and from that, it can be inferred that the strength of the relationship is lessened. The final Top Associated “combined” metric is a proposed combination of the first three top metrics with an additional ratio term which penalizes relations with high “far” distance co-occurrence counts with respect to their same sentence and near ones and is discussed in detail section 4.4.5. 4.4.4 ARTICLE STATISTICS TEMPORAL VIEW In addition to the ranked listings, we also see an interactive time line of the distribution of the article dates retrieved. Figure 4.9 shows what happens when a month is clicked on. Additionally, the months can be navigated via the “prior” and “next” links in the gray popup. This view allows us to visually assess whether there are periods of interest based on peaks in articles retrieved and the landing page “date range” filters can be used to focus on this time period only. Future work of interest could include trying to see whether communities found in the extended view visualization correlate with “events”, i.e. periods of time with increased article output. 18 Fig. 4.9 “Show Stats” Screen for Texas Representative Eddie Rodriguez 4.4.5 “COMBINED” METRIC & OTHER DISTANCE METRICS SUMMARY COMPARISON VIEW The final “Summary combined results” tab along the top of the window that pops up when “SHOW STATS” is clicked on from the landing page is of particular interest in that it allows a user to compare the ranked lists that each metric returns with respect to a particular entity type. The proposed “Combined” metric that appears in the final column is defined as: weight = (same sent + .5 * near sent + .1 same article) * boosting where boosting = combined co-occurrences / same article only co-occurences To clarify, “same sent” refers to the number in column 1, “near sent” to column 2 – column 1, “same article” to column 4, and “combined co-occurrences” to column 3. The coefficients associated with penalizing “near sentences” and “same article” co-occurrences (.5 and .1 respectively above) could and should be improved by having a person with domain knowledge, a political scientist specializing in Texas Politics in this instance, view the ranked results for each entity type of various, different politician graphs and then reorder the results of each if necessary thereby assessing their accuracy. It would then be a relatively straightforward process to use these newly labeled rankings to update the coefficients to produce results that more reflect the opinions of domain experts. Fig. 4.10 “Show Stats” Screen for Texas Representative Eddie Rodriguez 19 This process in itself is still subjective and could theoretically lead to over fitting if the training set is too small or is very different in behavior than the test data. The current coefficients were chosen by intuition to reflect near sentence occurrences being half as important as same sentence matches, and same article co-occurrences being a tenth as important. Its probably more likely that near sentence co-occurrences are more important and should be given a higher number, perhaps .75 and that same article ones should be slightly less important, perhaps a coefficient of .08, but that is left for future work. These results could be improved by additionally taking into account the number of sentences and unique entities in which co-occurrences are detected. For instance, in articles that have large numbers of entities relative to the number of sentences, and where it is possible that the article in itself is a listing of election results or candidates running for an office, it may be better to exclude “same article” relationships formed from the document since they will be largely noise. As a convenience method for showing what sort of rankings would entail from using different coefficient values, the visualization allows for user to pass them in, via the URL parameters “near_co” and “same_art_co”. For instance, the url: “explorer-view.html?s=Eddie Rodriguez&near_co=0.75&same_art_co=0.08” would show the Individual Star view for Eddie Rodriguez with the “Combined” metric using .75 and .08 for the new coefficients of near and same article only co-occurrences. Figure 4.10a shows the top revised results for Organizations most associated with Eddie Rodriguez using the above coefficients. Fig. 4.10a “Show Stats” Screen for Texas Representative Eddie Rodriguez with different coefficients When a user scrolls over or selects an entity from any of the lists in the “Summary comparison” tab, that entity is additionally highlighted in the other columns. For instance, Figure 4.10 shows the ranked position of the Workers Defense Project, an Organization entity type, among the difference metrics. As we can see, the Workers Defense Project currently does not appear among the upper ranked Organization entities for the third metric, which again is what the landing page visualization defaults to using, but we do see that it ranks highly according to our new metric since it has 2 same sentence occurrences, 4 – 2 = 2 near sentence occurrences, and relatively few distant occurrences (which is not visible in the figure, but is 1). Thus our new metric is computed as: weight = ( 2 + .5*2 + .1*1 ) * (2 + 2 + 1) / (1) = 3.1 * 5 = 15.5 If we then click on the Workers Defense Project (see APPENDIX F for full results) we see that in fact there is a strong direct relationship between the two as shown by this quote from an article from Texas Tribune1 in October 31, 2011. 1 http://www.texastribune.org/2011/10/31/cities-work-make-wage-theft-prosecution-priority/ 20 “ In Austin, the Workers Defense Project, a workplace justice group, is collaborating with state Representative Eddie Rodriguez, D-Austin, who sponsored the bill in the House, to set up a meeting to talk with Austin's police chief and the district and county attorneys about making wage theft an enforcement priority” 4.5 EXTENDED VIEW WITH COMMUNITY DETECTION When a user clicks on “Larger Network” for a given politician in the Table of Contents page, the data for the extended view network created during step 3.6 is displayed in an interactive webpage. This view is of the undirected graph with edges weighted according to the “Combined” metric described in section 4.4.5 that was created during the processing of the politician being studied. While viewing the graph, it is important to keep in mind that the edges are weighted by this “Combined” metric and do not just reflect the straight co-occurrence of two entities. This graph is much larger than the prior individual star view that could have at most N-1 edges displaying, where N is the number of entity nodes. This full graph on the other hand could possibly have N*(N-1)/2 edges if it is fully connected, i.e. if all nodes have connections with all other nodes. For this reason, in order to derive meaningful insight into the network it is necessary to be able to search for “communities” amongst the nodes and to filter out edges based on the weight, i.e. “importance”, of an edge between two entities. The idea of detecting communities in a network is similar conceptually to that of clustering in multivariate analysis and machine learning, and refers to a densely connected group of nodes that are well separated from the rest of the network. More formally and commonly, the definition of a community entails that the number of intra-community edges amongst the nodes of a single community be greater than the number of inter-community edges. There are a vast number of detection algorithms1 that can be used, but for our case we decided to go with the Louvain method[14] since it works on weighted graphs, provides a hierarchy of clusters, and is fast to run even on large graphs. We leverage an open-source JavaScript implementation of the method and D3.js to produce visualizations such as that of Figure 4.11 that shows the network formed by articles obtained and processed for Eddie Rodriguez. The graph in its entirety has 6226 unique entities/nodes and 572,965 edges. Our tool does community detection automatically, but allows the user to pass in a parameter to specify the number of communities they want the system to find and display, similar in fashion to selecting the “k” number of clusters to discover in k-means clustering. Selecting fewer communities to find and display leads to higher computational and time cost, and vise versa, more communities to find is less computational and time intensive. Additionally, the modularity of the overall communities is calculated and presented in the upper left hand side of the visualization. Modularity is a function that given a partition of nodes, tells you how good the community structure of that partition is [36] This value, between 0 and 1, measures the strength of division of the communities of the network and as such can be used as a sort of quality measure for the communities. Specifically, it is a weighted sum over all of the communities where the total number of inter-community edges is calculated for each node and then the expected number of edges under a random graph setting is subtracted from it. For more information on modularity and a concise overview of community detection in graphs in general see [35,37]. The tool also provides the ability to pass in a “threshold” parameter that removes edges whose weight is less than the threshold from the graph, and after which any nodes that no longer have edges are also filtered. In this case, setting a low weight threshold parameter is more computationally and time intensive, and results in more edges and nodes to visually portray which can make the 1 “Community structure in networks” by Arias & Ferrer-i-Cancho provides a good overview of community detection concepts & different methods. http://www.cs.upc.edu/~csn/lab/session5.pdf 21 visualization very congested. Conversely it does allow for the teasing out of more detailed relationships. In Figure 4.11 the community number parameter was set to 25 and the weight threshold was set to 15, which reduced the size of the graph dramatically to 952 nodes and 1869 edges. Community assignments are listed in the left hand side, and each community section can be expanded or collapsed, and the listing of nodes within each community can be selected therein or from the graph directly. Once a node is selected, its contents and edges are shown in the right hand side area. In Figure 4.11, the node for Eddie Rodriguez has been selected and dragged to a position for greater visibility. The right hand side shows he has 86 edges with weight greater than the threshold of 15 passed in, and the nodes associated with those edges are listed beneath him in descending order based on weight. The entity for “Robert Dulin” of type PERSON for instance has the highest “Combined” metric weight of 248.4, which means that he is the entity associated with Eddie Rodriguez and is thus listed at the top. The list of entities can also be ranked by “community” with the nodes within each community then ranked by weight or simply in alphabetical order by first name. Fig. 4.11 “Extended View” with Communities Screen for Eddie Rodriguez with Nodes Weighed By Strength. 4.5.1 ENTITY-ENTITY GRAPH EXPANSION VIEW Additionally, when a node is selected, the list of associated nodes in the right hand column can be selected in which case a reduced view of the graph is displayed showing the link between that node and the first node selected along with any other nodes they both have edges with. Figure 4.12 shows what happens when the entity “urban” of type MISC is selected from the prior results set for Eddie Rodriguez. 22 Fig. 4.12 “Extended View” with Eddie Rodriguez and then “urban” selected from right menu. In it, we can see that there are 4 other nodes associated with Eddie and “urban”; the locations “Springdale Farm”, and “East Austin”, and the politicians “David Simpson” and “Lois Kolkhorst”. As it turns Eddie Rodriguez, David Simpson, and Lois Kolkhorst are all part of the “Farm to Table Caucus” which appears a few entities down in the right hand side list, and that “Springdale Farm” is an “urban” farm located in the “urban” neighborhood of “East Austin”. This relationship is made even clearer if one selects the “urban” node directly from the graph as seen in Figure 4.13. The only new entities in this grouping are the newspaper that reported the most on these entities, the Texas “Tribune” of entity type ORGANIZATION, and Paula Foore, one of the owners of the “Springdale Farm”. This information was discovered simply by clicking on “Farm To Table” and “urban” in the prior Star view as seen in Fig 4.14. Fig. 4.13 “Extended View” with “urban” selected 23 Fig. 4.14 “Individual Star” Views of Eddie Rodriguez with “Farm To Table” & “urban” respectively selected To emphasize the importance of the community number and threshold parameters that a user may pass into the tool, Figure 4.15 shows the results of the prior example when the threshold parameter is passed in with a value 3 as opposed to 15, which allows for many more edges and nodes to remain. This time when the “urban” entity is selected we see many more entities than we did before including the Farm to Table Caucus itself, and the other owner of the Springdale Farm, Glen Foore. Fig. 4.15 “Individual Star” Views of Eddie Rodriguez with threshold parameter lowered to 3 Additionally, in order to facilitate searching for specific entities by name, the “search entity by name” input box at the top of the tool begins to auto-complete and show entities once three characters have been entered. Figure 4.16 shows the results when “Far” has been entered into the search box. Upon selecting one of the results, the node is displayed on the main graph, its community is opened in the left hand side “Communities Area”, and its information and edges are displayed in the right panel. 24 4.6 MEDIA ANALYSIS TOOLS In addition to the aforementioned tools that help users analyze politicians specifically from the vantage point of the people, bills, locations, and organizations with which they are associated in news articles produced from a set of news sources, the following tools allow the study of the news sources themselves. In our case study, the Austin American Statesman, the Dallas Morning News, the Houston Chronicle, the New York Times, the Texas Observer, and Texas Tribune were the sources used to find articles for the politicians studied. 4.6.1 MEDIA ANALYSIS TABLE VIEWS Figure 4.26 shows the top results obtained after all of the articles discovered for the 247 politicians amongst our news sources were processed. This table shows for which politicians the system obtained the most articles and how those articles are distributed amongst our 6 news sources. For instance, we see the Greg Abbott, the current Republican Governor of Texas, had 4200 articles processed and of those 1558 were from the Austin American Statesman, 88 from the Dallas Morning News, 780 from the Houston Chronicle, 449 from the New York Times, 81 from the Texas Observer and 1224 from the Texas Tribune. The next in the list include both of Texas’ federal senators, the Texas Speaker of the House, and the Lieutenant Governor of the State which is to be expected. Similarly Figure 4.47 shows the politicians for whom the least number of articles were processed, and they are composed of mostly junior State Representatives from largely rural districts and a judge elected in 2014 to the Texas Court of Criminal Appeals; none have appeared in the New York Times. Fig. 4.26 Media Analysis Table Sorted By Articles Successfully Processed. Top 5 Showing. Fig. 4.27 Media Analysis Table Sorted By Articles Successfully Processed. Bottom 5 Showing. 31 In addition to statistics on the breakdown of articles that were successfully processed by the system, the table shown in Figure 4.28 also continues to the right and shows how many articles were “skipped” by each source (shown in light red), and then goes further to show the breakdown of why they were skipped in the remaining five colored sections. These sections show articles returned from the internal site search engine for each news source that were not processed because they: • were sports articles (green), • were a duplicate result that was already processed in the same session ( tan ), • returned a reference to a link with no associated text possibly due to a subscription paywall which does not allow access to the article without a paid subscription or registration (purple), • returned an article that did not contain an exact reference to the politician used in the search query which could also be due to paywalls though not necessarily ( marsh green ) • were determined to be a list, possibly a candidate listing, high school scholarship winners listing, etc that were skipped to avoid introducing noise into the system. Fig. 4.28 Media Analysis Table Continued Generally speaking, articles skipped for the New York Times and Austin American Statesmen were due to paywalls. Articles skipped for the Dallas Morning News, Texas Observer, and Texas Tribune were due to the politician not being explicitly found in the article text returned. It should be noted that the internal search engines for both the Dallas Morning News and Texas Observer return a maximum of 100 articles per query, and as such are less represented as a whole. Finally articles skipped for the Houston Chronicle were mostly due to duplicate results being returned by the site’s internal search engine though it also contained many sports articles that were skipped; the Austin American Statesman, New Dallas Morning News and New York Times additionally contained sports articles that were skipped, but to a much lesser extent. Another view that allows for a more concise grouping of the top politicians as covered by the news sources in our case study is provided in Figure 4.29. In addition to providing a side-by-side ranked list, this “News Source Centric” view allows the user to hover over a politician to easily see how that person ranks in other news sources. In the figure it is easy to see how State Senator John Whitmire, whose district is in Houston, fairs among the different sources in relative article accounts. “Relative” here means that it is important to look at actual article counts since one table’s article counts may be low compared with another. For instance, there were more processed John Whitmire articles for the Texas Tribune than the Dallas Morning News. Abbreviations used in the level and position columns are provided at the bottom of the tool and all columns are sortable. 32 Fig. 4.29 Media Analysis News Source Centric Table View 4.6.2 MEDIA ANALYSIS HEAT MAPS OF TEXAS HOUSE, SENATE AND FEDERAL DISTRICTS In addition to the table view, there are heat maps showing the number of articles processed from different news sources for the districts pertaining to Texas House of Representatives, Texas Senate, and the Federal House. Figure 4.30 shows the map for the Texas House and in it the Texas Fig. 4.30 Heat Map of Articles Published by the Texas Observer on the Texas House Representatives by District Observer has been selected. The sources that can be selected are in the top upper right area next to a heat scale which means districts with the least articles published about their representative are 33 colored light yellow and those districts with the most articles are dark red. It is important to note that these heating color values are relative to the maximum articles published for a Texas Representative by the Texas Observer and not the aggregate whole of articles published at the State Representative level. The districts whose Representatives had the least number and the most number of articles published about them are automatically shown immediately below the sources. In this case Tom Craddick of District 82 had the most articles with 897, visible immediately below his picture, and District 58 with Representative J.D. Sheffield (not shown) had the least articles, zero in the case though to be fair there was a two way between himself and Dewayne Burns. It should be remembered though that there are 150 Texas House Districts as opposed to 31 Texas Senate Districts and as such the amount of media coverage for them as a whole is generally less than it is for Texas, and similarly Federal Senators. Figure 4.31 shows the Heat Map of articles published by the Dallas Mornings on the Texas Senate. We can see that the districts pertaining to and near the Dallas metropolitan area in the north eastern area of the state are more well represented than districts farther away from it. Fig. 4.31 Heat Map of Articles Published by the Dallas Morning News on the Texas Senate by District Figure 4.32 shows the Heat Map of articles published by the Houston Chronicle pertaining to Texas members of the Federal U.S. House of Representatives. In it we can again validate the importance of locality in terms of media coverage as the area around Houston is darkly shaded, and the Federal Representative with the most articles published about them, Republican John Culberson with 829 articles, represents a district within Houston. In both Figure 4.31 and 4.30 we can see that the area around the state capitol of Austin in central Texas has considerable shading as expected. Additionally there is an “Aggregate Results” and an “Aggregate Results Scaled” link among the list of sources that allows the user to view the news source article distributions as a collective whole, giving an overall average perspective of media coverage for the Texas House, Senate and Federal House. The Aggregate Results is a lump sum of all the sources per district, whereas the scaled version takes into to account the differences in total amounts of articles obtained for each source. For instance, for the Federal House Representatives map, the sums of the articles produced by each news source overall is as follows: AAS= 1962, DMN= 1751, HC= 13012, NYT= 1191, TXOB= 363 & TXTR= 3111. 34 Fig. 4.32 Heat Map of Articles Published by the Houston Chronicle on the US House of Representatives by District From these numbers we can see that the articles and subsequently the districts covered by the Houston Chronicle (HC) will dominate the aggregate lump view since 13012 articles pertaining to Federal Representatives were found there while the next highest count comes for the Texas Tribune with 3111 articles. Thus a possible solution for this disparity is to scale everything down to the least represented news source, the Texas Observer in the case, or conversely scale everything up to the most represented one, the Houston Chronicle, which is what the system does. Figure 4.33 shows the aggregate lump and aggregate scaled results for the Federal House of Representatives side by side. In it we can see some of the mass has been moved from the areas near Austin and Houston in the left “lump sum” version to areas near Dallas in the right “scaled” version. Fig. 4.33 Aggregate Lump (left) and Aggregate Scaled (right) maps of the US House of Representatives by District 35 See APPENDICES J, K, and L for all of the possible news source – congressional body combinations. As a final technical note, it should be clarified that the Media tools presented were not entirely automatically generated whereas the Network Visualizations for each Politician are. The only thing necessary to automatically utilize the media tools is simply a small script to handle the merging of a JSON file containing the final “politician processing” system statistics along with another JSON file containing the politician’s meta data; their district, party affiliation, image URL, etc. This procedure generates the data file used by the tools, and was handled here via a small R script. 5. Additional Analyses 5.1 AUTOMATED SUMMARIZATION OF POLITICIANS In an effort to get an overall picture of the political landscape of Texas given the data we obtained for our list of politicians, two approaches, one network centric and one based on text analysis through information retrieval and topic modeling techniques were considered. The former approach involves merging the “Extended” graphs for all the 247 politicians in our case study, and then running traditional network analysis on the large merged graph to determine the importance of politicians, organizations and other entities via various centrality measures. Although this method would inarguably produce interesting insight into the political landscape as a whole, it would do little in terms of a providing a simple summary of the issues and topics surrounding each politician. To clarify, the network-based approach would show the important connections between the different entities of our graph, but as it is still entity-based, since it was derived from an entity co-occurrence based metric, it would not explicitly extract the issues central to a politician. This information could be inferred of course from the graph by seeing for instance that a politician was highly connected to a particular organization that advocated a particular issue or by seeing a politician was connected closely to a given bill involving an issue. A negative consequence of this approach however is that if an entity from an article was not recognized correctly by our NER solution during processing, a connection to it will not be established. This fault also exists largely independent of how many articles that entity appears in since it is expected that if the NER did not identify properly in one article, it is probable that it doesn’t in another. NER tools can be pre-trained of course so this can be improved. However, and more importantly, if an article contains just the thoughts of a politician on a given issue, but contains no explicit mentions to an organization, politician, location, bill or other entity, that information will be lost in the graph. For this reason and for time considerations, this merging network based approach is left for future work. The second approach considered is based on textual analysis of the articles in which a politician occurs. By considering the combined text of the articles a politician occurs in as a single corpus, we can create a document term frequency matrix where each row is a particular article and the columns represent the terms (ie, single words, bigrams, trigrams) that occur in the corpus. We can then use this to calculate the TF-IDF (term frequency – inverse document frequency) of the corpus. This value gives a measure of how often a word appears through out the corpus with respect to the number of distinct articles it appears in. We can then use this TF-IDF value to filter out terms that appear all the time and provide little information or inversely those that occur vary rarely and could be noise. The art of deciding what stays in and what is filtered is very application specific and in this case was assessed by calculating a cutoff point based on a range of how many terms we desired appear overall in the filtered corpus. Then, using this refined corpus, we can run Latent Dirichlet Allocation [LDA 38] to uncover the “topics” (ie, issues, latent concepts) expressed within the articles for a single politician. The core assumption of LDA is that words in documents are generated by a mixture of topics with each topic itself being a distribution of words. In this algorithm, the number of topics is fixed initially and after it is run, each article in the corpus of a politician is assigned a topic. The topics themselves are composed of words that are ranked by those words with a higher probability of occurring. Because we have how many documents were assigned a given topic, we have the most frequent topics and can use these as a general summary of the issues surrounding a politician. 5 36 We ran LDA using 20 topics over each politician’s articles set for single words, bi-grams (pairs of words) and tri-grams separately, and gathered the results in a searchable tool shown in Figures 5.0 and 5.1. The number of topics to begin with is imprecise and based on seeing how the results performed for different topic values for a handful of cases normally used in the literature (5, 10, 20, etc.) The work in [16] is particularly interesting because it discovers the number of topics to use without an initial user given one. In the end, the optimal value is specific to each politician in question, and has a lot to do with the analysis one wants to present. As long as a sufficient number of topics were used, it had a less overall effect than the thresholds used in determining which terms to keep in the refined corpus. Figure 5.0 shows the results of searching for the topic “farms” Fig. 5.0 Searching for Farms in Texas in single term case In the figure, we can see a list of the politicians most relevant to the search query in question and see that Eddie Rodriguez has the highest relevance score. Upon selecting his name we see the all of the 20 topics found for Eddie Rodriguez along with the top 10 words for each, as shown in Figure 5.1a. The relevance score here is simply the percentage of articles for a politician that were assigned a topic with a term matching the query. In the figure we see the third topic, with relevance score 6.46 and highlighted in light red, matched the query and contains the words “food, markets, farms, regulations, milk” which all align with our prior examples showing Eddie’s involvement in agriculture and farmer’s markets issues. The relevance score means that 6.46% of the articles Representative Rodriguez occurred in were assigned this topic. Fig. 5.1a Finding a Politician whose articles in the news generally pertain to a specific topic using single words The above results were obtained using the “single words” version of tool that pertains to the LDA run using single word (i.e. 1-grams). Figure 5.1b shows the results using the bi-gram version of the tool. In it we see that “farms” is no longer in any of the topics, but the very first result which 15.65 of the documents involving Eddie Rodriguez are labeled with, includes the terms “Farmers Markets, Springdale Farm, Raw Milk, Cottage Food”. Whereas the first version is good at pulling out “broad” single-word themes that are each generally cohesive as a concept, the bi-gram version extracts proper nouns, with topics that are less cohesive/interpretable, and as such should be regarded as method of creating a word cloud of issues around someone. Both methods still are quite noisy, highly sensitive 37 to the parameters used when filtering, and the process of deciphering topics, i.e. “labeling” them, is a manual process that often requires prior knowledge of the domain space, Texas politics in our case. Because we have the article URLS, and their associated metadata, that the topics are assigned to, we can present them to the user to assist in the understanding and labeling of topics (not shown). Fig. 5.1b Finding a Politician whose articles in the news generally pertain to a specific topic using bi-grams The tri-gram topics (not shown) were more noisy, difficult to interpret and computationally more costly than the other two, providing little value outside of identification of 3-word proper nouns. Possible better approaches include trying to combine the three versions into one or instead focus on “phrase mining” techniques such as those shown in [17]. 5.2 AUTOMATED SUMMARIZATION OF COMMUNITIES If we return to Figures 4.24 and 4.25 presented in Section 4.5.4, we can see the details pertaining to a specific community detected within the context of the articles retrieved and processed for Eddie Rodriguez. The information provided in the figures provides details into who the central figures are within that community and additionally, the articles that are most prevalent within it. It would be useful however to have a way of automatically providing a description of the community at a higher level in order to give a more easily digestible global perspective of it. In that way we can then label all communities and allow the end user an additional perspective into the summaries as a whole. One way to do that is by treating the articles of a given community collectively as a single corpus. We can then analyze the corpus using the same procedure we use to “summarize politicians” described in the prior section; namely an initial TF-IDF procedure to filter terms and reduce noise, followed by performing Latent Dirichlet Allocation to derive topics. The main difference here though is that since we know how many entities from a community occurred in each article found in the community, we can weigh these articles by their relative importance. For instance, in Figure 4.25, the article1 including the most entities from the community will be weighed by a factor of 6. Additionally, we consider those articles with only one entity from the community as noise and exclude them from the corpus. As a proof of concept in Table 5.1 we show the results of applying this technique to the community from Figures 4.24 and 4.25 using single word, and the initial amount of topics being set to 5. The initial topics number is set low because if the number of communities is set Topic: 25.45% tax, strayhorn, rates, students, craddick, car, tesla, gambling, industry, cars Topic: 21.82% food, farmers, maps, markets, caucus, redistricting, doggett, plaintiffs, latino, map Topic: 18.79% craddick, gambling, interest, lenders, loans, loan, rates, tax, incentives, annual Topic: 18.79% energy, program, line, latin, market, sanchez, craddick, jobs, fashion, foreign Topic: 15.15% utility, uber, energy, tesla, rates, dealers, lyft, shoes, electric, stores Table 5.1 Summarizing a Community Via Topic Modeling with topics composed of single terms 1 http://www.texastribune.org/2013/05/30/bipartisan-caucus-lays-groundwork-food-movement/ 38 sufficiently high, we would hope that each community would encapsulate at most a few topics though this again varies per community and is based on the number of entities, articles, expansion/conductance values and the overall modularity of the communities We observe that due to the relatively low number of topics there is some overlap of concepts as highlighted in the second “single-words” topic in the above table, where the blue words refer to “agriculture” terms (farmers market, farm to table caucus) and the red refer to “redistricting” terms associated with articles discussing a lawsuit involving the re-drawing of district maps that would change U.S. Rep. Lloyd Doggett of Austin’s district to include Latino neighborhoods of San Antonio. Additionally for this proof of concept, we are not taking advantage of the stochastic nature of the results returned by LDA, which are in fact the likelihoods of a document belonging to any particular topic. The results above use just the most likely topic for assignment of documents, and as such lose the additional information provided in the posterior values of the model. This information will change the results quite a bit if a sizeable number of the articles have more than one topic with high probability. To illustrate the point, in the above example case, the community contains 170 articles and we calculate the difference between the most probable topic and the next most for each document to produce the histogram in Figure 5.2. Here we see that most of the documents have a clear topic assignment, but this may not always be the case. Future work will account for it. Fig. 5.2 Topic Assignment Verification for Example Case 5.3 MEDIA CENTRIC RESULTS OF CASE STUDY The distribution of articles that were downloaded, successfully processed and skipped considering all the news sources and politicians collectively are presented in Figure 5.3. In it we see that on average 800 articles were downloaded for the 247 politicians we are studying, and of those about 370 were processed successfully on average per politician while 429 were skipped for reasons explained in Section 4.6.0. These numbers illustrate the volatile nature of using unstructured text data for analysis and why preprocessing, filtering, and verifying data quality is so vital to systems such as our own. In order to delve more into where successfully processed articles where obtained, Figure 5.4 shows the distribution of processed articles by news source. In it, we see the relatively low numbers of articles coming from the Texas Observer and Dallas Morning News as previously mentioned, but we also see some outliers occurring in the Austin American Statesman, Texas Tribune and New York Times which pertain to Governor Abbott and US Senator Ted Cruz. Fig. 5.3 Distributions of Articles for Case Study 39 To go further into the average and more appropriately the median use cases for the sources, since none of the distributions pictured in Figure 5.4 are remotely normally distributed, we look at some summary statistics for each of the news sources in Figure 5.5. Fig. 5.4 Distributions of Articles Processed By News Source Figure 5.5 doesn’t provide much in terms of new insight except interestingly that the median articles processed per user are higher in the Dallas Morning News than expected Fig. 5.5 Summary Statistics By News Source (in order: AAS, DMN, HC, NYT, TXOB, TXTRB) In order to crudely determine if there is any bias in terms of reporting on Republicans versus Democrats as a whole among the different bodies we are studying we look at Figures 5.4 and 5.5. Fig. 5.4 Articles Processed For Texas House Representatives, Texas Senators, and the Texas Congress as a whole 40 "totalrels": 3282958, "totalinsts": 2653, "totalents": 43997, "BILL": 6355, "LOCATION": 7279, "ORGANIZATION": 7385 “MISC": 6843, "PERSON": 8064, "politician": 0, "Austin American Statesman": { "skipped": 0, "success": 104 }, "DALLAS MORNING NEWS": { "skipped": 0, "success": 52 }, "HOUSTON CHRONICLE": { "skipped": 0, "success": 959 }, "New York Times": { "skipped": 0, "success": 21 }, "TEXAS OBSERVER": { "skipped": 0, "success": 37 }, "Texas Tribune": { "skipped": 0, "success": 244 }, }, { "http://www.chron.com//news/houston-texas/article/Governor-decides-he-ll-fill-rest-of-seats-on-TSU-1611870.php": { "insts": 1, "processingtime": 1.0470820000000458, "rels": 33, "ents": 6, "result": "completed", "sents": 22 }, …….. "1": { "firstpass": [ { "articles": 15, "completed": 15 "processingtime": 73.8045259999999, "totalsents": 635, "totalrels": 44276, "totalinsts": 29, "totalents": 771, 47 "BILL": 75, “LOCATION": 75, "MISC": 75, "ORGANIZATION": 75, "politician": 0, "PERSON": 75, "HOUSTON CHRONICLE": { "skipped": 0, "success": 15 }, }, { "http://www.chron.com//default/article/Sesi-n-legislativa-especial-culmina-con-nota-2078478.php": { "insts": 2, "processingtime": 7.31278999999995, "rels": 3562, "ents": 56, "result": "completed", "sents": 42 }, … } APPENDIX F: WORKERS DEFENSE PROJECT – EDDIE RODRIGUEZ text results 48 APPENDIX H: INDIVIDUAL POLITICIAN RELATIVE ARTICLE VALUES. TEXAS & US SENATE, SELECT TEXAS ELECTED OFFICIALS APPENDIX H: INDIVIDUAL POLITICIAN SCALED ARTICLE VALUES. TEXAS & US SENATE, SELECT TEXAS ELECTED OFFICIALS 49 APPENDIX I: INDIVIDUAL POLITICIAN RELATIVE ARTICLE VALUES. TEXAS HOUSE OF REPRESENTATIVES 2015 APPENDIX I: INDIVIDUAL POLITICIAN SCALED ARTICLE VALUES. TEXAS HOUSE OF REPRESENTATIVES 2015 50 APPENDIX J: TEXAS SENATE, TEXAS HOUSE & FEDERAL ARTICLE DISTRIBUTIONS BY DISTRICT FOR AUSTIN AMERICAN STATESMAN, DALLAS MORNING NEWS, AND HOUSTON CHRONICLE. COLUMNS ARE LEGISLATIVE BODIES AND ROWS ARE NEWS SOURCES 51 APPENDIX K: TEXAS SENATE, TEXAS HOUSE & FEDERAL ARTICLE DISTRIBUTIONS BY DISTRICT FOR THE NEW YORK TIMES, TEXAS OBSERVER AND TEXAS TRIBUNE. COLUMNS ARE LEGISLATIVE BODIES AND ROWS ARE NEWS SOURCES 52 APPENDIX L: TEXAS SENATE, TEXAS HOUSE & FEDERAL ARTICLE DISTRIBUTIONS BY DISTRICT FOR AGGREGATE RESULTS AND AGGREGATE RESULTS SCALED. COLUMNS ARE LEGISLATIVE BODIES AND ROWS ARE NEWS SOURCES 53