scieee AI-readable full text Open interactive document viewer

Cohort size and labour-market outcomes

Roth, Duncan,Moffat, John D.,Garloff, Alfred

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Roth, Duncan; Moffat, John D.; Garloff, Alfred Book Cohort size and labour-market outcomes IAB-Bibliothek, No. 367 Provided in Cooperation with: Institute for Employment Research (IAB) Suggested Citation: Roth, Duncan; Moffat, John D.; Garloff, Alfred (2018) : Cohort size and labourmarket outcomes, IAB-Bibliothek, No. 367, ISBN 978-3-7639-4121-6, W. Bertelsmann Verlag (wbv), Bielefeld, https://doi.org/10.3278/300969w This Version is available at: https://hdl.handle.net/10419/280220 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-sa/4.0/ Duncan Roth 367 Cohort size and labour-market outcomes 367 Cohort size and labour-market outcomes Duncan Roth Inaugural-Dissertation zur Erlangung der wirtschaftswissenschaftlichen Doktorwürde des Fachbereichs Wirtschaftswissenschaften der Philipps-Universität Marburg eingereicht von Duncan Roth M. Sc. aus Darmstadt Erstgutachter: Prof. Dr. Bernd Hayo Zweitgutachter: Prof. Dr. Michael Kirk Einreichungstermin: 10. Oktober 2016 Prüfungstermin: 20. Dezember 2016 Erscheinungsort: Marburg Hochschulkennziffer: 1180 Herausgeber der Reihe IAB-Bibliothek: Institut für Arbeitsmarktund Berufsforschung der Bundes agentur für Arbeit (IAB), Regensburger Straße 100, 90478 Nürnberg, Telefon (09 11) 179-0 Redaktion: Martina Dorsch, Institut für Arbeitsmarktund Berufsforschung der Bundesagentur für Arbeit, Telefon (09 11) 179-32 06, E-Mail: [email protected] Gesamt herstellung: W. Bertelsmann Verlag, Bielefeld (wbv.de) Rechte: Kein Teil dieses Werkes darf ohne vorherige Genehmigung des IAB in irgendeiner Form (unter Verwendung elek tro nischer Systeme oder als Ausdruck, Fotokopie oder Nutzung eines anderen Vervielfältigungsverfahrens) über den persönlichen Gebrauch hinaus verarbeitet oder verbreitet werden. © 2018 Institut für Arbeitsmarktund Berufsforschung, Nürnberg/ W. Bertelsmann Verlag GmbH & Co. KG, Bielefeld In der „IAB-Bibliothek“ werden umfangreiche Einzelarbeiten aus dem IAB oder im Auftrag des IAB oder der BA durchgeführte Untersuchungen veröffentlicht. Beiträge, die mit dem Namen des Verfassers gekenn zeichnet sind, geben nicht unbedingt die Meinung des IAB bzw. der Bundesagentur für Arbeit wieder. ISBN 978-3-7639-4120-9 (Print) ISBN 978-3-7639-4121-6 (E-Book) ISSN: 1865-4096 Best.-Nr. 300969 www.iabshop.de www.iab.de Bibliografische Information der Deutschen Nationalbibliothek Die Deutsche Nationalbibliothek verzeichnet diese Publikation in der Deutschen Nationalbibliografie; detaillierte bibliografische Daten sind im Internet über http://dnb.ddb.de abrufbar. Herausgeber der Reihe IAB-Bibliothek: Institut für Arbeitsmarktund Berufsforschung der Bundes - agentur für Arbeit (IAB), Regensburger Straße 100, 90478 Nürnberg, Telefon (09 11) 179-0 Redaktion: Martina Dorsch, Institut für Arbeitsmarktund Berufsforschung der Bundesagentur für Arbeit, 90327 Nürnberg, Telefon (09 11) 179-32 06, E-Mail: [email protected] Titelfoto: © gettyimages/Alexander Spatari Gesamtherstellung: wbv Media GmbH & Co. KG, Bielefeld (www. wbv.de) Rechte: Kein Teil dieses Werkes darf ohne vorherige Genehmigung des IAB in irgendeiner Form (unter Verwendung elektro nischer Systeme oder als Ausdruck, Fotokopie oder Nutzung eines anderen Vervielfältigungs verfahrens) über den persönlichen Gebrauch hinaus verarbeitet oder verbreitet werden. © 2018 Institut für Arbeitsmarktund Berufsforschung, Nürnberg/ wbv Publikation, ein Geschäftsbereich der wbv Media GmbH & Co. KG, Bielefeld In der „IAB-Bibliothek“ werden umfangreiche Einzelarbeiten aus dem IAB oder im Auftrag des IAB oder der BA durchgeführte Untersuchungen veröffentlicht. Beiträge, die mit dem Namen des Verfassers gekennzeichnet sind, geben nicht unbedingt die Meinung des IAB bzw. der Bundesagentur für Arbeit wieder. ISBN 978-3-7639-4126-1 (Print) ISBN 978-3-7639-4127-8 (E-Book) ISSN 1865-4096 DOI 10.3278/300985w Best.-Nr. 300985 www.iabshop.de www.iab.de Bibliografische Information der Deutschen Nationalbibliothek Die Deutsche Nationalbibliothek verzeichnet diese Publikation in der Deutschen Nationalbibliografie; detaillierte bibliografische Daten sind im Internet über http://dnb.ddb.de abrufbar. Dieses E-Book ist auf dem Grünen Weg Open Access erschienen. Es ist lizenziert unter der CC-BY-SA-Lizenz. 3IAB-Bibliothek 367 Inhalt Acknowledgements ................................................................................. 5 Einleitung – German Summary ............................................................... 7 Problem statement, structure and contribution of the dissertation .... 17 John Moffat and Duncan Roth The cohort size-wage relationship in Europe ....................................... 25 (published LABOUR: Review of Labour Economics and Industrial Relations, 2016, 30 (4): 415–32) Abstract ..................................................................................................................................... 25 1 Introduction ................................................................................................................. 25 2 Estimation .................................................................................................................... 28 2.1 Data ................................................................................................................................ 28 2.2 Empirical model .......................................................................................................... 30 2.3 Identification ............................................................................................................... 33 3 Results ........................................................................................................................... 35 4 Conclusion .................................................................................................................... 38 Acknowledgements ............................................................................................................... 39 References ................................................................................................................................ 39 Appendix ................................................................................................................................... 42 Supplementary material ...................................................................................................... 46 References ................................................................................................................................ 59 Alfred Garloff and Duncan Roth Regional population structure and young workers’ wages ................... 61 (forthcoming in U. Blien, K. Kourtit, P. Nijkamp and R. Stough (eds): Modelling Aging and Migration Effects on Spatial Labor Markets, Springer) Abstract ..................................................................................................................................... 61 1 Introduction ................................................................................................................. 61 2 Population structure and wages ............................................................................ 64 3 Youth-population structure in Western Germany ............................................. 66 4 Empirical analysis ....................................................................................................... 69 4.1 Data ................................................................................................................................ 69 4.2 Sample and descriptive statistics .......................................................................... 70 4.3 Empirical model and identification ....................................................................... 73 5 Results ........................................................................................................................... 75 6 Conclusion .................................................................................................................... 83 Acknowledgements ............................................................................................................... 84 References ................................................................................................................................ 84 Appendix ................................................................................................................................... 88 Supplementary material ...................................................................................................... 89 References ................................................................................................................................ 105 Inhalt IAB-Bibliothek 367 4 John Moffat and Duncan Roth Cohort size and youth labour-market outcomes: the role of measurement error ................................................................ 107 (forthcoming in Economics Bulletin) Abstract ..................................................................................................................................... 107 1 Introduction .................................................................................................................. 107 2 Literature review ......................................................................................................... 109 3 Empirical analysis ....................................................................................................... 111 3.1 Data ................................................................................................................................ 111 3.2 Variables and sample ................................................................................................ 113 3.3 Model ............................................................................................................................. 117 4 Results ........................................................................................................................... 119 5 Conclusion .................................................................................................................... 123 Acknowledgements ............................................................................................................... 124 References ................................................................................................................................ 124 Appendix ................................................................................................................................... 126 Supplementary material ...................................................................................................... 134 References ................................................................................................................................ 164 Duncan Roth Cohort size and transitions into the labour market .............................. 165 Abstract ..................................................................................................................................... 165 1 Introduction .................................................................................................................. 165 2 Literature and hypotheses ........................................................................................ 167 3 Empirical analysis ....................................................................................................... 170 3.1 Data ................................................................................................................................ 170 3.2 Sample and variables ................................................................................................. 171 3.3 Model ............................................................................................................................. 176 4 Results ........................................................................................................................... 178 4.1 Baseline results ........................................................................................................... 178 4.2 Discussion of the hypotheses ................................................................................. 182 4.3 Alternative explanations .......................................................................................... 183 4.4 Inclusion of individuals with zero search duration .......................................... 185 5 Conclusion .................................................................................................................... 186 Acknowledgements ............................................................................................................... 188 References ................................................................................................................................ 188 Appendix ................................................................................................................................... 191 Supplementary material ...................................................................................................... 192 References ................................................................................................................................ 203 Abstract .................................................................................................... 205 Kurzfassung .............................................................................................. 207 5IAB-Bibliothek 367 Acknowledgements This book contains the doctoral thesis that I wrote at Philipps-Universität Marburg and the Institute for Employment Research. For his continued support, comments and suggestions I am particularly grateful to my supervisor, Prof Bernd Hayo. I would also like to thank Prof Michael Kirk for agreeing to supervise my work as well as for his feedback. Further, I would like to thank Prof Tim Friehe for acting as chair of the defence committee. I am also grateful to Stefan Fuchs for supporting me and giving me the opportunity to complete this thesis. Furthermore, I would like to acknowledge the exchange that I had with a number of different people whose input contributed to my work: my co-authors John and Alfred, my colleagues at Marburg and at IAB as well as the participants of various workshops and conferences. These acknowledgements would be incomplete without reference to my family and to Julia to whom this book is dedicated. Duncan Roth Marburg, February 2018 7IAB-Bibliothek 367 Einleitung – German Summary Diese Arbeit setzt sich aus vier separaten Essays zusammen, die den Zusammenhang zwischen regionalen Bevölkerungsstrukturen und verschiedenen Arbeitsmarktergebnissen zum Thema haben. The cohort size-wage relationship in Europe Das erste Papier mit dem Titel The cohort size-wage relationship in Europe untersucht den Zusammenhang zwischen der Größe einer Gruppe, deren Mitglieder eine ähnliche Berufserfahrung (oder ein ähnliches Alter) und ein vergleichbares Ausbildungsniveau aufweisen, auf die Löhne, die von den Mitgliedern einer solchen „Kohorte“ realisiert werden. Basierend auf der Annahme, dass Personen innerhalb einer Kohorte substituierbar sind, dies über verschiedene Kohorten hinweg aber nur unvollständig möglich ist, lässt die ökonomische Theorie vermuten, dass Änderungen in der Größe einer Kohorte zunächst deren Grenzproduktivität beeinträchtigt. Auf Wettbewerbsmärkten sollte dies eine Anpassung in den kohortenspezifischen Löhnen verursachen. Im Fall abnehmender Grenzproduktivität lässt sich dieser Zusammenhang genauer spezifizieren: Ceteris paribus, sollte ein Anstieg in der Größe einer Kohorte dazu führen, dass die Grenzproduktivität innerhalb der Kohorte und dadurch auch die erzielten Löhne sinken. Theoretische Modelle legen darüber hinaus nahe, dass ein vergleichbarer Mechanismus auch im Fall unvollkommenen Wettbewerbs greift, wenn Löhne durch Verhandlungen zwischen Arbeitgeberund Arbeitnehmervertretern festgesetzt werden. In der bestehenden empirischen Literatur wird mehrheitlich ein negativer Lohneffekt nachgewiesen. Darüber hinaus gibt es Hinweise, dass die Größe dieses Effekts mit dem Ausbildungsniveau der Kohorte ansteigt. Eine Schwierigkeit, den Lohneffekt empirisch zu bestimmen, besteht darin, dass nicht davon ausgegangen werden kann, dass die Zugehörigkeit einer Person zu einer bestimmten Kohorte zufällig ist. Vielmehr ist in Betracht zu ziehen, dass Personen durch eigene Entscheidungen ihre Kohortenzugehörigkeit beeinflussen können. Im Fall einer durch Berufserfahrung (oder Alter) und Ausbildungsniveau bestimmten Kohorte kann dies einerseits dadurch geschehen, dass Personen in Regionen migrieren, die für die Höhe der von ihnen erzielten Löhne förderlich sind. Andererseits bestimmen Ausbildungsentscheidungen darüber, welcher Kohorte eine Person angehören wird. Beide Mechanismen verwandeln die Kohortengröße selbst in eine endogene Variable, sodass die Anwendung des Kleinste-Quadrate-Schätzers möglicherweise verzerrte Ergebnisse liefert. Der Beitrag dieses Papiers besteht darin, eine Identifikationsstrategie zu verwenden, die in der Lage ist, beide Ursachen der Endogenität zu berücksichtigen, während Einleitung – German Summary IAB-Bibliothek 367 14 Zeitpunkt jüngere Altersgruppen typischerweise kleiner sind als ältere und somit auch das Arbeitsangebot – gemessen an der Zahl der Person – geringer ausfallen sollte. Gleichzeitig sollte insbesondere der Anteil derer, die dem Arbeitsmarkt nicht zur Verfügung stehen, aufgrund verstärkter Teilnahme an Bildungsmaßnahmen höher ausfallen. Eine negative Korrelation zwischen der Höhe des kohortenspezifischen Arbeitsangebots und der Höhe des Messfehlers sollte sich dann einstellen, wenn der Unterschied im Anteil der Nichtteilnehmer die Unterschiede in der Größe der Altersgruppen überwiegt. Diese Beziehung wird auch durch die Instrumentierung nicht aufgelöst, da Kohorten, die in der Gegenwart relativ klein sind, auch zu einem früheren Zeitpunkt vergleichsweise klein gewesen sein sollten. Bei älteren Gruppen sollte diese Art des Messfehlers eine geringere Rolle spielen, da der Anteil der Arbeitsmarktteilnehmer deutlich höher ausfallen sollte. Für die Schätzung des Effekts von Kohortengröße auf Arbeitslosigkeit und Beschäftigung sind junge Altersgruppen daher weniger geeignet. Da bestehende Studien oftmals jüngere Altersgruppen in die empirische Analyse aufgenommen haben, ist dieses Ergebnis für die Literatur relevant, da es die Frage aufwirft, in welchem Maß die bisherigen Ergebnisse von Messfehlern in der Kohortenvariable beeinträchtigt sind. Cohort size and transitions into the labour market Das letzte Papier befasst sich mit dem Zusammenhang zwischen der Kohortengröße beim Eintritt in den Arbeitsmarkt und der Dauer bis zum Beginn der ersten Beschäftigung. Da die Auswirkungen auf die Suchdauer bisher noch nicht untersucht worden sind, leistet dieses Papier zum einen durch die Wahl einer neuen Ergebnisvariable einen Beitrag zur Literatur. Zum anderen unterscheidet es sich von den zuvor besprochenen Papieren dadurch, dass hier nicht der kontemporäre Zusammenhang zwischen der Größe einer Kohorte und einem bestimmten Arbeitsmarktergebnis betrachtet wird. Stattdessen geht es um die Auswirkung, die die Kohortengröße zu einem bestimmten Zeitpunkt – nämlich beim Eintritt in den Arbeitsmarkt – auf nachfolgende Entwicklungen, in diesem Fall die Suche nach Beschäftigung, hat. Aufgrund dieser Änderung im zeitlichen Kontext des untersuchten Zusammenhangs weist das Papier auch einen Bezug zu einer weiteren Literatur auf, in der der Einfluss von Konjunktureffekten beim Arbeitsmarkteintritt auf zukünftige Arbeitsmarktergebnisse untersucht wird. Da in diesen Analysen andere Eintrittsbedingungen – z. B. die Größe der Eintrittskohorte – unberücksichtigt bleiben, können die Ergebnisse dieses Papiers auch für diese Literatur von Bedeutung sein. Um Hypothesen zu bilden, wie sich die Größe der Eintrittskohorte auf die anschließende Dauer der Suche nach Beschäftigung auswirkt, wird auf die Literatur zum bereits im Kontext des vorigen Papiers beschriebenen Zusammenhang Einleitung – German Summary 15IAB-Bibliothek 367 zwischen Kohortengröße und Arbeitslosigkeit zurückgegriffen. Demnach wäre es zunächst möglich, dass in größeren Eintrittskohorten aufgrund der stärker ausgeprägten Konkurrenz auf dem Arbeitsmarkt länger gesucht werden muss, bevor eine Beschäftigung gefunden werden kann. Dieser Effekt könnte jedoch dadurch abgeschwächt (oder umgekehrt) werden, dass Personen, die den Arbeitsmarkt als Teil einer großen Kohorte betreten, Beschäftigungen aufnehmen, die unter ihrem Anforderungsprofil liegen. Schließlich besteht die Möglichkeit, dass es in größeren Eintrittskohorten zu kürzeren Suchdauern kommt, wenn Unternehmen angesichts eines gestiegenen Arbeitsangebots junger Altersgruppen Stellen schaffen. Grundlage für die Untersuchung des beschriebenen Zusammenhangs bilden Daten zu Absolventen von Ausbildungsprogrammen. Dieser Fokus ist in mehrerer Hinsicht sinnvoll: Erstens ist mit den vorliegenden Daten eine Identifizierung des Orts und des Zeitpunkts des Ausbildungsabschlusses sowie des ersten nachfolgenden Beschäftigungsverhältnisses möglich (für andere Gruppen, z. B. die Hochschulabsolventen, liegen vergleichbare Angaben zum Studienabschluss nicht vor). Zweitens, beinhaltet diese Gruppe nicht nur Personen ähnlichen Alters, sondern auch einer vergleichbaren beruflichen Qualifikation. Im Gegensatz zu ausschließlich nach Alter abgegrenzten Kohorten sollte in diesem Fall, in dem die Eintrittskohorte auf dem Erwerb eines berufsqualifizierenden Abschlusses beruht, auch die Relevanz der Kohorte für den Arbeitsmarkt höher sein, was das im vorigen Papier beschriebene Problem des Messfehlers aufgrund fehlender Teilnahme am Arbeitsmarkt reduzieren sollte. Schließlich ist die Gruppe der Auszubildenden an sich relevant, da es sich hierbei um einen in Deutschland verbreiteten Weg handelt, mittels dessen junge Personen den Arbeitsmarkt betreten. Durch diese Einschränkung sind die Ergebnisse jedoch nicht zwangsläufig auf andere Gruppen, wie die der Hochschulabsolventen oder der Geringqualifizierten übertragbar, für die sich der untersuchte Zusammenhang womöglich anders dargestellt hätte. In der empirischen Analyse werden zwei Datenquellen aus Deutschland verwendet, die bereits im Kontext des zweiten Papiers beschrieben worden sind: die Integrierten Erwerbsbiografien (IEB) sowie die Stichprobe der Integrierten Arbeitsmarktbiografien (SIAB). In einem ersten Schritt muss die zentrale erklärende Variable – die Größe der Eintrittskohorte – geschätzt werden, indem auf Grundlage der IEB die Zahl der Personen innerhalb eines bestimmen Zeitraums und in einer bestimmten Arbeitsmarktregion berechnet wird, die eine Reihe an Bedingungen erfüllen, sodass sie als Absolventen eines Ausbildungsprogramms gezählt werden können. Die der eigentlichen Regressionsanalyse zugrundeliegende Stichprobe wird hingegen aus SIAB-Daten gewonnen. Berücksichtigt werden männliche Personen, die zwischen Januar 1999 und Oktober 2012 im Alter von 19 bis 23 Jahren eine Ausbildung abgeschlossen haben. Um zu vermeiden, dass Ab- Einleitung – German Summary IAB-Bibliothek 367 16 solventen aus früheren Jahren systematisch längere Suchdauern aufweisen, werden unterschiedliche Analysen für verschiedene Zeiträume durchgeführt, über die alle Individuen in der Stichprobe ab dem Zeitpunkt des Ausbildungsabschlusses beobachtet werden (3 Monate, 6 Monate, 1 Jahr, 2 Jahre). Erfolgt innerhalb eines solchen Zeitraums ein Übergang in Beschäftigung, so wird er als solcher gezählt, wohingegen für Personen, deren Übergänge zu einem späteren Zeitpunkt erfolgen, die Information genutzt wird, dass es innerhalb des Beobachtungszeitraums nicht zu einer Beschäftigungsaufnahme gekommen ist. Da es sich bei der zu erklärenden Variable um eine Dauer handelt, werden für die empirische Untersuchung Methoden der Verweildaueranalyse und insbesondere das Cox-Modell genutzt. Die Ergebnisse legen nahe, dass Absolventen, die als Teil einer größeren Kohorte in den Arbeitsmarkt eintreten, schneller eine Beschäftigung finden. Allerdings zeigt sich, dass dieser Effekt nur dann signifikant ist, wenn der dreimonatige Beobachtungszeitraum angewendet wird; bei längeren Zeiträumen ist der Effekt hingegen kleiner und statistisch insignifikant. Für den sechsmonatigen Beobachtungszeitraum stellen sich jedoch sehr ähnliche Ergebnisse ein, sobald nicht nur für die Größe der Kohorte beim eigenen Eintritt in den Arbeitsmarkt kontrolliert wird, sondern auch die Größe der nachfolgenden Eintrittskohorte berücksichtigt wird. Diese Ergebnisse liefern somit keine Evidenz für die erste Hypothese, dass Mitglieder größerer Eintrittskohorten aufgrund verstärkter Konkurrenz längere Suchdauern haben. Weitere Untersuchungen zeigen, dass die Größe der Kohorte keinen negativen Effekt auf die Höhe der Löhne hat, die im ersten Beschäftigungsverhältnis nach der Ausbildung erzielt werden, und auch nicht zu einer höheren Wahrscheinlichkeit führt, dass eine andere als eine sozialversicherungspflichtige Art der Beschäftigung – z. B. eine geringfügige Beschäftigung – aufgenommen wird. Diese Ergebnisse sprechen somit auch gegen die zweite Hypothese, dass sich kürzere Suchdauern bei größeren Eintrittskohorten durch eine Selektion in weniger anspruchsvolle Beschäftigungen erklären lassen. Alternative Erklärungen für die empirischen Befunde – Selektion der Absolventen in Regionen mit kürzeren Suchdauern nach Beendigung der Ausbildung oder Unterschiede in der Zusammensetzung größerer Kohorten hinsichtlich der Produktivität ihrer Mitglieder – werden ebenfalls nicht durch die empirische Evidenz gestützt. Abschließend finden sich auch keine Belege dafür, dass die Ergebnisse auf die Tatsache zurückzuführen sind, dass das Cox-Modell Personen, die keine Suchdauer aufweisen, da sie direkt nach Beendigung der Ausbildung eine Beschäftigung finden, nicht berücksichtigen kann. Wenn die Ergebnisse auch keinen direkten Beleg für die dritte Hypothese darstellen, dass Unternehmen angesichts größerer Eintrittskohorten neue Stellen schaffen, so ist diese Erklärung doch mit dem Befund kompatibel, dass es Mitgliedern größerer Kohorten schneller gelingt, nach Beendigung der Ausbildung eine Beschäftigung zu finden. 17IAB-Bibliothek 367 Problem statement, structure and contribution of the dissertation The aim of this thesis is to contribute to the understanding of how changes in cohort size affect various labour-market outcomes. It is therefore related to a large body of literature that has developed out of the desire to shed light on the implications of the large post-World-War-II birth cohorts entering the US labour markets from the late 1960s onwards (Freeman, 1979; Welch, 1979) and that has since continued to address the relationship between population structure and the labour market. In the part of this literature that is most relevant to my work the subject of interest is typically constituted by the effect that the size of an age group has on group-specific outcomes which is motivated by the assumption that members of different age groups are only imperfectly substitutable and as such compete for jobs mainly within their group. This assumption in turn reflects the view that differently aged individuals can be expected to differ with respect to the amount of work experience and human capital that they have acquired (Welch, 1979) and as long as human capital is a determinant of a worker’s productivity on the job, there should be limits to the extent to which substitution across age groups is possible. In terms of economic models this assumption is reflected in workers of different age groups representing separate factors of production (Berger, 1983; Connelly, 1986; Card and Lemieux, 2001). The central explanatory variable in this context is based on the concept of a cohort, which measures the size of a specific age group. The extant literature differs with respect to exactly how a cohort is defined, with the underlying age groups being either relatively broad, often representing the size of the youth population (Korenman and Neumark, 2000; Shimer, 2001; Biagi and Lucifora, 2008), or being based on single-year age groups (Wright, 1991; Mosca, 2009; Brunello, 2010). Other studies have employed specifications in which cohort size is delineated according to years of experience rather than age (Welch, 1979), with the former variable being argued to be more relevant to determining whether individuals are substitutable. Furthermore, the cohort that an individual belongs to may not only be determined by his age or experience, but also by his level of education (Welch, 1979; Wright, 1991; Mosca, 2009; Brunello, 2010). Such a specification allows for the effects of cohort size to differ between different levels of education but also imposes the assumption that differently educated individuals are active on separate labour markets. The most commonly used outcome variables in this literature and the ones most relevant to this thesis are cohort-specific wages as well as employment and unemployment rates. In the case of a perfectly competitive labour market an Problem statement, structure and contribution of the dissertation IAB-Bibliothek 367 18 increase in cohort size should lead, ceteris paribus, to a fall in the wages earned in that age group if there is diminishing marginal productivity in production – an illustration of the effects of an outward shift in the labour-supply curve. This relation is shown formally by Brunello (2010), while Michaelis and Debus (2011) develop a model of imperfectly competitive labour markets in which wages are determined by bargaining between firms and monopoly unions. They show that in most cases an increase in the size of an age group will decrease the wages of that group. According to Stapleton and Young’s (1988) diminishing-substitutability hypothesis the negative relationship between cohort size and wages should be more pronounced among the highly educated as the former are less easily substitutable across age groups. The majority of the available empirical research provides evidence for a negative wage effect of cohort size and often finds results to be in line with the diminishing-substitutability hypothesis (Welch, 1979; Wright, 1991; Brunello, 2010). The possibility that wages might not fully adjust in response to changes in cohort size provides the possibility of a relationship between cohort size and cohort-specific employment or unemployment rates. Fertig and Schmidt (2004) argue that larger cohorts may have a higher degree of bargaining power which may help to prevent a downward wage adjustment, while a fixed number of jobs for a specific age group also constitutes a reason for changes in cohort size translating into (un-)employment adjustments (Korenman and Neumark, 2000). In contrast to the case of wage outcomes, there is no consensus on the sign of this relationship. A number of empirical analyses have yielded evidence that increases in cohort size lead to a larger group-specific (Korenman and Neumark, 2000; Biagi and Lucifora, 2008) or overall unemployment rate (Garloff et al., 2013) which would appear to suggest that there are negative labour-market consequences of belonging to a larger cohort. These findings, however, contrast with an argument proposed by Shimer (2001) – which he also supports with empirical evidence – that regions in which the share of young age groups is larger should experience lower youth and overall unemployment rates. This hypothesis rests on the assumption that an increase in the share of youths, who are often either unemployed or poorly matched and thus willing to take up or to switch jobs, makes it easier for firms to fill vacancies, so that an anticipated increase in the youth share is met by an expansion in the number of jobs offered. Skans (2005) also provides evidence that supports the hypothesis that the youth unemployment rate falls with the size of the youth cohort. The above literature forms the basis for this thesis. The first three papers address issues which in my view represent shortcomings in the available research on cohort-size effects and provide empirical evidence to support this view. In Problem statement, structure and contribution of the dissertation 19IAB-Bibliothek 367 contrast, the fourth paper analyses the effect on an outcome variable that has so far not been the subject of research in this literature, and treats cohort size as a labour-market entry condition rather than a contemporaneous explanatory variable. The core of each paper is formed by an empirical analysis that assesses the effects of cohort size on individual-specific or group-specific outcomes. Moreover, each paper comes with supplementary material which further elaborates on arguments made in the corresponding paper and provides the results of various sensitivity analyses. The topic of the first paper is the effect of cohort size on wages and how the former varies across educational groups. It argues that the identification strategy that has so far been used in studies on the wage effect is not suited to purge the endogeneity of the cohort-size variable that can arise because of selected migration into high-wage areas. While the limited amount of cross-national migration makes disregarding this possibility appear innocuous when the size of the cohort is measured at the country level, the former becomes much more of a concern at the regional level. Moreover, in light of what cohort size is supposed to measure – the supply of labour from a specified group within a labour market – it would appear more appropriate to base this variable on regions since they are likely to closer resemble the delineation of labour markets than countries. The results provide evidence – at least for the largest educational group – that the proposed identification strategy produces qualitatively different results – a negative significant effect as opposed to an insignificant one – compared to the previously employed identification strategy. Identifying the effects of interest is complicated by the fact that the size of a cohort arguably cannot be treated as an exogenous variable: individuals are not randomly allocated to certain cohorts, but can influence which group they belong to at a given point in time through decisions pertaining to migration and investment in education. Since experimental data is not available, this paper – as well as the two subsequent ones – employs an instrumental-variables strategy in order to arrive at a consistent estimate of the cohort-size effect. This approach is not without problems of its own, though. Since two-stage least squares (2SLS) estimation is less efficient than ordinary least squares (OLS), the effect of interest is estimated less precisely. Moreover, its interpretation depends on the chosen instrument which in this case is given by the size of the cohort observed a certain number of years earlier when the members of the cohort were younger by the same amount of years. The estimated wage effect therefore stems from a change in (contemporaneous) cohort size that is caused by a change in its lagged value. This could be problematic if, for example, those who later on migrate represent a selected group of individuals. Finally, the instrument itself, while displaying a high degree of (partial) correlation with the endogenous cohort-size variable, might be Problem statement, structure and contribution of the dissertation IAB-Bibliothek 367 20 put into question since the problem of selected migration may simply be shifted from the individual to his parents. The contribution of the second paper is twofold. First, it aims to produce insights into the mechanisms that are behind the negative wage effect of cohort size and finds that a substantial part of this effect is due to selection into lowerpaying occupations and, to a lesser extent, industries. Second, it raises the question to what extent the cohort-size variables that are used in other studies contain measurement error. If this variable, as discussed above, is supposed to measure group-specific supply within an actual labour market, it is questionable whether the typically employed administrative units represent a reasonable basis as their delineations are not designed to produce entities within which a specified group of individuals competes for employment. Since (random) measurement error in an explanatory variable leads to attenuation bias, it is possible that the magnitude of the wage effect has been underestimated in previous studies (including the former). The paper proceeds by estimating two separate models in which the cohort-size variable is either derived from administrative units or from the functional labour-market regions derived by Eckey et al. (2006). The former model produces smaller cohort-size coefficients, thereby providing evidence that the choice of the underlying spatial entity is relevant in terms of the magnitude of the estimated effects. The second paper also differs from the first with respect to the data it uses, which in this case come from register entries rather than from a survey, which may provide more reliable information about certain variables such as wages. Moreover, the data come from a single country, Germany, rather than from a sample of European countries – a feature which might be attractive in terms of reducing the potential of confounding influences. When data from different countries (or regions) is pooled in order to estimate a given model, the implicit assumption is made that the relationship is the same in each case, though the inclusion of appropriate fixed effects allows for countryor region-specific intercepts. However, differences in national labour-market institutions, for example, could lead to the relationship between cohort size and the outcome variable being structurally different between countries. Since the institutional framework can be expected to be more homogenous within a country, use of data from a single country arguably reduces this problem. Estimating the effect of changes in cohort size on the (un-)employment rate within that cohort is the subject of the third paper. In light of the conflicting empirical evidence that has been produced by the extant literature, this paper provides new insights into this relationship. The main motivation for this analysis, however, is the hypothesis that cohort-size variables are subject to measurement Problem statement, structure and contribution of the dissertation 21IAB-Bibliothek 367 error when they contain very young age groups, which is often the case in the existing literature. Since a substantial share of individuals in these groups will not be available to the labour market – primarily, though not exclusively, due to participation in education – an age-specific cohort-size variable will provide only a poor measure of labour supply in that group, which in turn may affect size and sign of the estimated effects. The paper develops an argument of nonclassical measurement error which the previously discussed identification strategy is unable to correct for. The results indeed show that the estimated cohort-size effects change drastically depending on the age range of the sample. The final paper addresses the relationship between cohort size and the transition into the labour market by analysing the former’s effect at the time of labour-market entry on the duration of search for employment. What sets this analysis apart from the other papers is not only that a new outcome variable is being analysed, but rather that cohort size is not treated as a contemporaneous explanatory variable. Instead the variable represents a condition under which entry into the labour market took place and which might affect subsequent outcomes. Given this setting, there are parallels between the subject of this paper and a recent literature that analyses the long-run effects of the state of the business cycle at the time of labour-market entry on future labour-market outcomes (Stevens, 2007; Kahn, 2010; Brunner and Kuhn, 2014; Cockx and Ghirelli, 2016). In my view, the contribution of this thesis to the existing cohort-size literature has been to raise questions about the adequacy of the existing empirical methodology to identify the effects of interest and to provide evidence that these matters can have a substantial impact on the results. My understanding from reading the literature is that the cohort-size variable is supposed to quantify the supply of labour by a specified group whose members are reasonably similar so that they can be regarded as substitutable in production and who are active on the same labour market. If this reading is correct, questions about measurement are bound to arise and two have been addressed in this thesis: is it important to base the cohort-size variable on spatial units that approximate actual labour markets and how does the inclusion of very young age groups, substantial shares of which are often not available to the labour market, affect the results. Moreover, conceptualising cohort size as a labour-market entry condition, raises questions for future research that aim at assessing the long-run consequences of having entered the labour market as part of a large or small cohort. Problem statement, structure and contribution of the dissertation IAB-Bibliothek 367 22 References Berger, M.C. (1983) Changes in Labor Force Composition and Male Earnings: A Production Approach, Journal of Human Resources, 18, 177–96. Biagi, F. and Lucifora, C. (2008) Demographic and education effects on unemployment in Europe, Labour Economics, 15, 1076–101. Brunello, G. (2010) The effects of cohort size on European earnings, Journal of Population Economics, 23, 273–90. Brunner, B. and Kuhn, A. (2014) The impact of labor market entry conditions on initial job assignments and wages, Journal of Population Economics, 27, 705–38. Card, D. and Lemieux, T. (2001) Can Falling Supply Explain the Rising Return to College for Younger Men? A Cohort-Based Analysis, Quarterly Journal of Economics, 116, 705–46. Cockx, B. and Ghirelli, C. (2016) Scars of recessions in a rigid labor market, Labour Economics, 41, 162–76. Connelly, R. (1986) A Framework for Analyzing the Impact of Cohort Size on Education and Labor Earning, Journal of Human Resources, 21, 543–62. Eckey, H.-F., Kosfeld, R. and Türck, M. (2006) Abgrenzung deutscher Arbeitsmarktregionen, Raumforschung und Raumordnung, 64, 299–309. Fertig, M. and Schmidt, C. (2004) Gerontocracy in Motion: European Cross-Country Evidence on the Labor Market Consequences of Population Ageing, in R. Wright (ed.) Scotland’s Demographic Challenge, Scottish Economic Policy Network, Stirling, Glasgow. Freeman, R.B. (1979) The effect of demographic factors on age-earnings profiles, Journal of Human Resources, 14, 289–318. Garloff, A., Pohl, C. and Schanne, N. (2013) Do small labor market entry cohorts reduce unemployment?, Demographic Research, 29, 379–406. Kahn, L.B. (2010) The long-term labor market consequences of graduating from college in a bad economy, Labour Economics, 17, 303–16. Korenman, S. and Neumark, D. (2000) Cohort Crowding and Youth Labor Markets: A Cross-National Analysis, in D.G. Blanchflower and R.B. Freeman (eds.) Youth Employment and Joblessness in Advanced Countries, University of Chicago Press, Chicago. Michaelis, J. and Debus, M. (2011) Wage and (un-)employment effects of an ageing workforce, Journal of Population Economics, 24, 1493–511. Mosca, I. (2009) Population Ageing and the Labour Market, Labour, 23, 371–95. Shimer, R. (2001) The Impact of Young Workers on the Aggregate Labor Market, The Quarterly Journal of Economics, 116, 969–1007. Problem statement, structure and contribution of the dissertation 23IAB-Bibliothek 367 Skans, O.N. (2005) Age effects in Swedish local labor markets, Economics Letters, 86, 419–26. Stapleton, D.C. and Young, D.J. (1988) Educational Attainment and Cohort Size, Journal of Labor Economics, 6, 330–61. Stevens, K. (2007) Adult Labour Market Outcomes: the Role of Economic Conditions at Entry into the Labour Market, available at http://www.iza.org/conference_ files/SUMS2007/stevens_k3362.pdf Welch, F. (1979) Effects of Cohort Size on Earnings: The Baby Boom Babies’ Financial Bust, Journal of Political Economy, 87, S65–S97. Wright, R.E. (1991) Cohort size and earnings in Great Britain, Journal of Population Economics, 4, 295–305. IAB-Bibliothek 367 30 The cohort size-wage relationship in Europe data from the year 2011 cannot be used and that only those individuals that are observed in adjacent years can be retained. In terms of countries our final sample includes observations from Austria, Belgium, Bulgaria, Cyprus, Czech Republic, Denmark, Estonia, France, Greece, Hungary, Italy, Latvia, Lithuania, Luxembourg, Malta, Norway, Poland, Romania, Slovakia, Spain and Sweden.4 For each of these countries EU-SILC provides information on an individual’s residence at the NUTS1 level. This piece of information is crucial as it allows construction of the cohortsize variable at the regional level. The countries listed above provide us with a total of 56 NUTS1 regions. 2.2 Empirical model The dependent variable of our model is given by the natural logarithm of the purchasing power parity (PPP)-adjusted hourly wage of individual i in experience group j, with educational qualification e, residing in region r at time t, wijert. This variable is constructed by first adjusting annual wage income for inflation using the GDP deflator (base year: 2010). This variable is then divided by the countryspecific PPP-factor from the base year, as provided by Eurostat (see Friedrich, 2015). This quantity is then divided by the number of hours usually worked per week, which are multiplied by the number of weeks per year and the fraction of the year spent working as reported by the individual. To reduce the risk of measurement error due to changes in the number of hours worked over the year, we restrict our sample to those individuals that have been working either exclusively full-time or exclusively part-time during the income-reference period. The main explanatory variable is the relative size of the experience cohort to which the individual belongs. This variable’s specification follows from the assumptions made about the group with whom the individual is substitutable. First, we follow the literature (Card and Lemieux, 2001; Brunello, 2010) in assuming that substitutability is possible within but not across educational categories. The level of education in EU-SILC is given by the 1997 system of the International Standard Classification of Education (ISCED-97) which allows for cross-country comparisons of educational qualifications. This variable assigns a value from 0 (pre-primary education) to 5 (first stage of tertiary education) to every individual. Because of top-coding, individuals with ISCED 6 (second stage of tertiary education) cannot be identified separately but are subsumed into category 5. We 4 Observations from the following countries are excluded: Ireland and the UK (income-reference period is not the preceding calendar year as it is for other countries); Germany, the Netherlands and Portugal (no information on region of residence); Croatia (due to unavailability of data, the instrumental variable cannot be constructed); Slovenia (information on the degree of urbanisation missing); Finland and Iceland (year of birth as well as all agerelated variables are not recorded precisely, presumably for disclosure reasons). 31 Estimation Chapter 1 follow Brunello (2010) in combining individual categories into larger educational groupings: ISCED 0–2 includes all individuals with at most lower secondary education, ISCED 3–4 combines upper secondary and post-secondary, non-tertiary education and ISCED 5 contains individuals with completed tertiary education. Second, we assume that individuals compete for jobs within regions rather than countries. This approach is preferable for two reasons. First, it allows the use of inter-regional variation in cohort size to identify the former’s effect on wages. More substantively, we argue that labour markets are more likely to exist at a sub-national level because of limitations to mobility or because information about job opportunities decreases with distance from an individual’s place of residence. Ideally, we would base cohort size on spatial entities which are delineated in a way that the working population residing in such an area would be exclusively employed there and vice-versa. But while such functional units have been designed for individual countries (see Eckey et al., 2006, for Germany), no comparable units have been defined for the European level. But the fact that functional labour markets tend to be found to be relatively small suggests that the use of NUTS1 regions as approximations of regional labour markets is preferable to the use of countries. Finally, we choose to define cohort size in terms of labour-market experience rather than age. Within an educational grouping, years of work experience provides a measure of the human capital that individuals have had a chance to accumulate on the job. The use of experience thus provides a better measure of substitutability in the labour market than age and also ties in with Welch’s (1979) proposed career-phase model in which workers with different levels of experience differ in terms of the tasks they can perform, making them only imperfectly substitutable. However, results comparable to those presented in Section 3 are obtained when an age-specific cohort-size variable is used.5 If individuals are not at all substitutable across experience groups, the appropriate cohort-size variable would be defined simply as the ratio of individuals of experience j with education e in region r at time t, Njert , relative to the number of all individuals with education e in region r at time t, Nert . But since it is likely that individuals are substitutable if they have similar but not necessarily the same level of experience, we follow Wright (1991) and Brunello (2010) in calculating the numerator of the cohort-size variable as a weighted average of the number of individuals with up to two years more or two years less work experience6: 5 The results of this and all subsequently mentioned robustness checks are available upon request. 6 Notice that the use of V-shaped weights implies that substitutability decreases with the difference in experience levels (see Wright, 1991, for a discussion). Comparable results to those presented in Section 3 are obtained when different specifications of the numerator are used. IAB-Bibliothek 367 32 The cohort size-wage relationship in Europe CSjert = (1/9)Nj – 2, ert + (2/9)Nj – 1, ert + (3/9)Njert + (2/9)Nj + 1, ert + (1/9)Nj + 2, ert Nert [1] Because official statistics regarding the size of education-experience groups at a regional level are not available, these quantities are estimated from the EU-SILC dataset using the adjusted sampling weights. For each of the three educational groups, the sample from which cohort size is calculated includes males and females who are either employed or unemployed. Given that our focus is on individuals in the early stages of their career, a large share of inactive individuals are in the process of acquiring education and including those observations would, for example, mean including all individuals enrolled in tertiary education in the construction of cohort size for ISCED grouping 3–4, which in turn would lead to an artificial jump in cohort size once these individuals have completed education and entered the ISCED 5 grouping. The inactive are therefore excluded from the construction of the cohort-size variable. However, comparable results are obtained when cohort size is constructed from all individuals regardless of their economic status. The sample from which the cohort-size variable is constructed is restricted to the working-age population (age groups 16–65) within each educational grouping. Cohort size is computed for up to 11, 9 and 5 years of experience for ISCED 0–2, ISCED 3–4 and ISCED 5, respectively. These upper limits are imposed for two reasons: first, our interest lies in individuals who are at an early stage of their career. Second and as discussed in Section 2.3, we want to ensure that the age groups which are used as instruments for the experience-based cohort-size variable do not contain individuals who are older than 15 in order to rule out issues of regional self-selection. Furthermore, the denominator includes individuals with up to 47, 45 and 41 years of experience in the case of ISCED 0–2, ISCED 3–4 and ISCED 5. These values are derived from assuming education-specific ages at entry into the labour market of 18, 20 and 24, respectively. In Section 2.3 we discuss how these assumptions fit the actual data from the regression sample. Since there is an upper and a lower limit to experience, the construction of cohort size has to be adjusted at the corners of this range by reallocating the weights that would otherwise have been attached to the experience groups outside the specified range. At the lower limit, cohort-size for experience groups 0 and 1 are constructed as follows (with corresponding constructions at the education-specific upper limits): CS0, ert = (6/9)N0, ert + (2/9)N1, ert + (1/9)N2, ert Nert [1a] 33 Estimation Chapter 1 CS1, ert = (3/9)N0, ert + (3/9)N1, ert + (2/9)N2, ert + (1/9)N3, ert Nert [1b] In terms of control variables, xijrt , we include a constant, individual-level regressors (indicators for working part-time, being married, the degree of urbanisation of the place of residence, and occupational indicator variables), experience-related regressors (experience and squared experience), region-specific regressors (region dummies), time-specific regressors (year dummies) and region-by-time regressors (the regional unemployment rate). Further details on these variables as well as descriptive statistics are given in Tables A1 and A2 in the Appendix. Our empirical model, which we estimate separately for each level of education (ISCED 0–2, ISCED 3–4, ISCED 5), is given by Equation 2 (throughout the remainder of the paper the e subscript is dropped): ln[wijrt ] = α CSjrt + β xijrt + uijrt [2] We exclude female observations from the estimation of Equation 2 to avoid the issue of selected labour-market participation. To account for the sampling design weighted regressions are performed. Finally, as the main regressor in our model is defined at a higher level of aggregation than the dependent variable, standard errors are clustered at the level of the region-experience cell (see Moulton, 1990). 2.3 Identification An obstacle to identifying the wage effect of cohort size using ordinary least squares (OLS) estimation is that individuals are not necessarily randomly allocated into cohorts. Rather membership of a specific cohort is potentially the result of individual self-selection into educational groups or regions. This would be the case if an individual’s expectation about future wages affected the decision to acquire a specific level of education (thereby affecting education-specific cohort size) or if labour-market prospects induced migration into a different region (thereby affecting region-specific cohort size). Due to the comparatively low costs of moving between regions (as opposed to countries) the second type of selfselection is of particular concern although freedom of movement of labour within the EU implies that migration between countries may also be significant. OLS is likely to underestimate the depressing effect of cohort size, if individuals select into educational groups or regions that are characterised by higher wages. To identify the effect of cohort size on wage consistently we therefore employ IV estimation. IAB-Bibliothek 367 34 The cohort size-wage relationship in Europe Recent contributions to the literature on cohort-size effects on wages either do not address the issue of endogeneity (Mosca, 2009) or acknowledge selfselection into educational groups, while implicitly disregarding self-selection through migration (Sapozhnikov and Triest, 2007; Brunello, 2010).7 The latter studies use contemporaneous age-specific cohort size as an instrument which is not differentiated by education. We argue that this approach suffers from the disadvantage of not addressing individual self-selection from migration. To assess this hypothesis we construct an instrumental variable (IV1) that corresponds to the one described above: IV1gkt = (1/9)N(g – 2)rt + (2/9)N(g – 1)rt + (3/9)Ngrt + (2/9)N(g + 1)rt + (1/9)N(g + 2)rt Nrt [3] Subscript g refers to age and the numerator is a weighted average of the number of individuals in a region that are two years younger, one year younger, the same age, one year older and two years older. Our preferred instrument (IV2) deals with both self-selection into educational groupings and self-selection into geographical areas. It is the relative size of the age group in the region that is fourteen years younger, fourteen years ago. Since the first year of sampling is the year 2004 and regional population data are available for most NUTS1 regions from the year 1990 onwards, fourteen years represents the longest feasible lag. Comparable instruments have been employed in the analysis of cohort-size and unemployment (Korenman and Neumark, 2000; Shimer, 2001; Garloff et al., 2013; Moffat and Roth, 2016). IV 2jrt = (1/9)Ng – 2 – 14, rt – 14 + (2/9)Ng – 1 – 14, rt – 14 + (3/9)Ng – 14, rt – 14 + (2/9)Ng + 1 – 14, rt – 14 + (1/9)Ng + 2 – 14, rt – 14 N rt – 14 [4] This variable is a natural predictor for our cohort-size variable as, in the absence of migration and natural changes in population, the individuals on which the instrument is based will be the same as those on which education-specific cohort-size is based, only that they are observed at different points in time. This association between the endogenous cohort-size variable and the instrument is supported by the first-stage test-statistics. In addition to not varying across education, both of the above instruments are defined in terms of age rather than experience. This requires us to specify a link between an individual’s age and years of experience. We do this with imputed 7 In contrast, Morin (2015) uses a natural experiment given by a reform to the educational system in a specific Canadian province to identify cohort-size effects. 35 Results Chapter 1 age values which are defined as the sum of assumed entry age (18, 20 and 24 for educational groupings ISCED 0–2, ISCED 3–4 and ISCED 5) and number of years of experience. We compare actual and imputed age and find that the distribution of true age is centred on the imputed age values in the majority of cases. We prefer matching cohort size and the instrument using an imputed age rather than actual age as this ensures that individuals in the same experience cohort are assigned the same value of the instrument, thereby avoiding identification of cohort-size effects from within-cohort variation in the instrument. To avoid the inclusion of an age group where individuals may make their own decisions about where to reside, the age groups in the instrument are restricted to be no older than 15. This implies an upper age limit of 29 for those in the sample and we therefore exclude observations from the regression who are older than 29 (though raising the limit to 32 does not affect the results). 3 Results Table 1 shows the coefficients of cohort size, experience and experience squared for each of the three education groups obtained by OLS, two-stage-least squares (2SLS) estimation using the instrument of Brunello (2010) and Wright (1991) – in the column headed IV1 – and 2SLS using an instrument based on lagged population sizes – in the column headed IV2. A full set of results can be found in Table A3 of the Appendix. Each of the three specifications produces negative cohort-size coefficients for all ISCED groups and, with the exception of ISCED 5, the coefficients of model IV2 are more negative than those of either OLS or IV1. This finding is in line with the previous discussion that due to their inability to account for self-selection through migration into high-wage areas the identification strategies of the latter models will underestimate the negative wage effect of cohort size. However, the size of the standard errors of specifications IV1 and IV2 suggest that the difference between the two point estimates is not statistically significant. In the case of ISCED 0–2 we find that none of the cohort-size coefficients is statistically significant. From a theoretical perspective (see Card and Lemieux, 2001; Brunello, 2010), this finding is compatible with individuals with different levels of experience in this educational category being easily substitutable for each other. Accordingly, the estimated effect of an increase in cohort-size of one standard deviation on an individual’s wage is comparatively small at -3% for specification IV2. In contrast, we find that cohort size has a considerable and statistically significant effect on the wages of individuals with completed secondary and post-secondary, non-tertiary education (ISCED 3–4): based on specification IV2, an increase in IAB-Bibliothek 367 36 The cohort size-wage relationship in Europe cohort size of one standard deviation decreases the wages of individuals in the affected cohort by 10%, ceteris paribus. The fact that the estimated effect is larger for ISCED 3–4 than for ISCED 0–2 is in line with Stapleton and Young’s (1988) ‘diminishing-substitutability hypothesis’ that differently aged/experienced workers are less easily substitutable at higher levels of education. Table 1: Cohort size coefficients obtained from weighted regression (OLS and 2SLS) ISCED 0–2 ISCED 3–4 ISCED 5 OLS IV1 IV2 OLS IV1 IV2 OLS IV1 IV2 Cohort Size -1.42 (1.11) -0.95 (3.22) -2.54 (2.67) -0.21 (0.94) -3.91 (3.50) -12.02** (4.70) -1.87 (1.56) -3.65 (7.44) -1.94 (7.66) Experience 0.08*** (0.01) 0.08*** (0.01) 0.08*** (0.01) 0.05*** (0.01) 0.06*** (0.01) 0.07*** (0.01) 0.10*** (0.03) 0.11** (0.05) 0.10** (0.05) Experience2-0.00*** (0.00) -0.00*** (0.00) -0.00*** (0.00) -0.00*** (0.00) -0.00*** (0.00) -0.00*** (0.00) -0.01* (0.00) -0.01 (0.01) -0.01 (0.01) N(inds) 7,364 7,364 7,364 19,785 19,785 19,785 4,973 4,973 4,973 N(cells) 2,180 2,180 2,180 2,499 2,499 2,499 1,338 1,338 1,338 N(clusters) 596 596 596 553 553 553 319 319 319 F-stat 115.81*** 124.82*** 53.24*** 44.63*** 27.37*** 21.43*** ME (std) -1.71% -1.14% -3.05% -0.17% -3.28% -10.07%** -2.28% -4.46% -2.37% Control variables from Table A1 are included. ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors are clustered at the region-experience level. N(inds): number of individual observations. N(cells): number of region-experience-year cells. N(clusters): number of clusters. F-stat refers to the first-stage F-statistic on the significance of the instrument in the first-stage regression of the endogenous cohort-size variable. ME(std) shows the percentage change in the hourly wage for a change in cohort size by one standard deviation. The results for ISCED 5 do not display a similar pattern: though all coefficients are negative, the point estimates of specification IV2 are smaller than the corresponding results for ISCED 3–4 as well as the coefficients from IV1 in the same educational group. Moreover, none of the coefficients on cohort size are statistically significant. The only other study of which we are aware which has found larger negative cohort-size effects for those with secondary education than those with tertiary education is Dooley (1986) who obtained this result using Canadian data. There are several potential reasons for this finding. From an empirical perspective, the comparatively small number of experience cells (5) reduces the variation from which the effect of cohort size can be identified (as evidenced by the number of region-year-experience or region-experience cells in the case of ISCED 5). The size of the standard errors in the IV estimations is also a consequence of the instrument’s decrease in predictive power with respect to cohort size (as evidenced by the comparatively small values of the F-statistic). From an economic 37 Results Chapter 1 perspective, the smaller size of the coefficients may be explained by greater segmentation of the labour market at higher levels of education. In other words, individuals with higher levels of education operate in more heterogeneous labour markets and are therefore less substitutable with individuals with the same level of education, irrespective of their level of experience. An alternative explanation is that individuals, once they have attained tertiary education, are more affected by the mechanism discussed by Berger (1989) that leads individuals in larger cohorts to obtain less human capital and therefore relatively high wages when young. This would be the case if, as seems likely due to opportunities to pursue postgraduate education or receive advanced on-the-job training, individuals with ISCED 5 have more scope for differentiated levels of human capital than individuals with ISCED 0–2 or ISCED 3–4. If those with ISCED 5 in large cohorts choose not to take these opportunities, there then will be a weaker relationship between cohort size and wages within this group. Another possibility is that the effect of cohort size occurs more through unemployment than wages for individuals with tertiary education. However, it is unclear why the wages of those with tertiary education would be more rigid or more influenced by unions so we regard the previous explanations for our inability to find significant and negative effects for this group as more credible. The first-stage F-statistics are above the rule-of-thumb value of 10 for each educational group, suggesting that there is no problem of weak instruments. The size of the test statistics decreases with the level of education which implies that lagged age-structures are a better predictor for education-specific cohort size of the less educated. A possible reason is that geographic mobility increases with the level of education. The experience variables show the standard positive but diminishing effect of experience on wages. If experience dummies are used, this pattern is also found and the estimated effects of cohort size are very similar to those reported above. The coefficients of the other control variables (reported in Table A3) are also in line with expectations. Specifically, higher regional unemployment is associated with lower wages while being married and living in a more urbanised environment have a positive effect on wages. The coefficients on the occupational dummies are also statistically significant and of the expected pattern. To put our results for ISCED 3–4 into perspective, we compare them with those of Brunello (2010), who uses a dataset and empirical model that is comparable to the one used in this paper. One difference between the two analyses is that his data is aggregated at the level of the country-year-age cell, while this analysis is based on individual-level variables. However, estimating Equation 2 after averaging all variables over the observations within a region-year-experience cell and weighting the regression by the number of observations per cell (adjusted for IAB-Bibliothek 367 38 The cohort size-wage relationship in Europe the sampling weights, see Angrist and Pischke, 2009), we obtain results that are very similar to those shown in Table 1. Another difference is Brunello’s (2010) use of a log-log specification. When we adopt this approach, we find that an increase in cohort size by 1% is predicted to decrease the mean wage of individuals in that cohort by 0.098% (IV1). This result is comparable to the predicted change of -0.069%, as estimated by Brunello (2010) for those aged below 35 in educational group ISCED 3–4. However, employing our preferred instrument (IV2) yields a predicted decrease of -0.288%, four times the size of the effect found by Brunello (2010). Notwithstanding the possibility that the difference is due to differences in the sampling period and range of countries included in the analysis, this supports our contention that, as discussed previously, the contemporaneous age-cohort size is unable to deal with self-selection through migration and use of this instrument leads to an underestimation of the true cohort-size effect. 4 Conclusion The aim of this paper has been the identification of the causal effect of cohort size on the wages of young men at the start of their career in Europe. Ex ante, the direction of this effect is unclear. If labour markets are perfectly competitive and differently aged workers are only imperfectly substitutable within each educational group, members of larger cohorts can be expected to receive lower wages as a result of their lower marginal productivity. However, in an environment of imperfect competition, increases in cohort size may have no or only a limited effect on wages if these are sufficiently rigid (in which case (un-)employment rates would be expected to change) or even a positive effect if larger cohorts are able to exert larger bargaining power. Identification of this effect is complicated by the fact that an individual’s cohort is likely to be the result of self-selection into educational groups and self-selection into geographic areas. Unlike earlier papers that have looked at this question using crosscountry European data, our approach addresses both types of self-selection by using the size of the population 14 years younger, 14 years ago as an instrument for cohort size. We also use regions rather than countries as the spatial unit since the former provides greater variation in the cohort-size variable and are also likely to provide a better approximation of actual labour markets than countries. Our results show that cohort size represents a significant and negative determinant of wages for young males with secondary but not for those with less than secondary or tertiary education. This suggests that the projected fall in the share of young people in the labour force will put upward pressure on the wages of those with secondary education – the largest group in the labour force. The finding that those with lower levels of education do not experience 39 References Chapter 1 a negative effect is consistent with the ‘diminishing-substitutability hypothesis’. We suggest that the failure to find a significant effect for those with tertiary education may be due to greater market segmentation among the more highly educated or greater scope for obtaining different levels of human capital after the completion of formal education. The effect of cohort size is not found to be statistically significant for any of the educational groupings if the IV strategy does not address the potential for self-selection through migration. This implies that the earlier work on the cohort size-wage relationship which did not address this source of endogeneity may have underestimated the true effect. Acknowledgements The authors would like to thank Eckhardt Bode, Bernd Hayo, Michael Kirk and Christian Traxler as well as the participants of the MACIE brownbag seminar, the MAGKS seminar, the 2012 Dutch Demographic Day, the 2013 Alpine Population Conference, the 2013 Congress of the European Association of Labour Economists (EALE), the 2013 Congress of the European Regional Science Association (ERSA) and the 8th European Workshop on Labour Markets and Demographic Change for valuable comments. This paper uses data from the European Union Statistics on Income and Living Conditions (EU-SILC). The results and conclusions are those of the authors and not those of Eurostat, the European Commission or any of the national statistical authorities whose data have been used. References Alsalam, N. (1985) The Effect of Cohort Size on Earnings: an Examination of Substitution Relationships, Working Paper 21, Centre for Economic Research, University of Rochester, Rochester. Angrist, J. and Pischke, J.-S. (2009) Mostly harmless econometrics: an empiricist’s companian, Princeton University Press, Princeton. Berger, M.C. (1983) Changes in Labor Force Composition and Male Earnings: A Production Approach, Journal of Human Resources, 18, 177–96. Berger, M.C. (1985) The Effect of Cohort Size on Earnings Growth: A Reexamination of the Evidence, Journal of Political Economy, 93, 561–73. Berger, M.C. (1989) Demographic Cycles, Cohort Size, an Earnings, Demography, 26, 311–21. Berger, M. and Schaffner, S. (2015) A Note on How to Realize the Full Potential of the EU-SILC Data, ZEW Discussion Paper 15-005, Centre for European Economic Research, Mannheim. IAB-Bibliothek 367 46 The cohort size-wage relationship in Europe Supplementary material The first part of this section provides further information on how the endogenous cohort-size variable, which is defined in terms of experience, is matched with the age-based instrumental variable. In the second part, we present the results of a variety of sensitivity analyses as evidence for the robustness of our findings. S1 Experience, age and imputed age In our empirical model, cohort size is constructed on the basis of experience, whereas the variable which the former is instrumented with is age-specific. This requires us to specify a relation between age and experience. We do this by imputing an age variable that is constant for all individuals within an experience group. This variable is defined as the sum of years of experience and the education-specific age at which entry into the labour market is assumed to take place. Specifically, for each of the three educational groupings imputed age is defined as follows: ISCED 0–2: age_imputed = 18 + years of experience (range: 18–29) ISCED 3–4: age_imputed = 20 + years of experience (range: 20–29) ISCED 5: age_imputed = 24 + years of experience (range: 24–29) The distribution of actual age for given values of imputed age is shown below for each educational group (Figures S1–S3). The graphs illustrate that actual age is indeed centred on the corresponding value of imputed age in the vast majority of cases. For higher values imputed age does not appear to lie in the centre of the actual age distribution. This, however, results from individuals whose actual age is above 29 being excluded from the sample. 47 Supplementary material Chapter 1 Figure S1: Age and imputed age (ISCED 0–2) 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 FrequencyFrequencyFrequencyFrequency FrequencyFrequencyFrequencyFrequency FrequencyFrequencyFrequencyFrequency Notes: The dark line indicates the value of imputed age from 18 (top-left) to 29 (bottom-right) Source: EU-SILC (authors’ calculations). Age Age 16 18 20 22 24 26 28 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 IAB-Bibliothek 367 48 The cohort size-wage relationship in Europe Figure S2: Age and imputed age (ISCED 3–4) Age 16 18 20 22 24 26 28 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 600 400 200 0 Frequency Frequency Frequency Frequency Frequency Frequency Frequency Frequency Frequency Frequency Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Notes: The dark line indicates the value of imputed age from 20 (top-left) to 29 (bottom-left) Source: EU-SILC (authors’ calculations). Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 49 Supplementary material Chapter 1 Figure S3: Age and imputed age (ISCED 5) Notes: The dark line indicates the value of imputed age from 24 (top-left) to 29 (bottom-right) Source: EU-SILC (authors’ calculations). 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 200 100 0 Frequency Frequency Frequency Frequency Frequency Frequency Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 Age 16 18 20 22 24 26 28 IAB-Bibliothek 367 50 The cohort size-wage relationship in Europe S2 Sensitivity analysis This section starts by illustrating the effect on the 2SLS cohort-size coefficients (corresponding to specification IV2) of excluding individual regions, years or experience groups from the sample. As can be seen from Figures S4–S8, the results are typically very close to the coefficient of the full model and always lie within the former’s 95% confidence interval. Figure S4: Excluding individual regions from the sample (ISCED 0–2) AT1 AT2 AT3 BE1 BE2 BE3 BG3 BG4 CY0 CZ0 DK0 EE0 ES1 ES2 ES3 ES4 ES5 ES6 ES7 FR1 FR2 FR3 FR4 FR5 FR6 FR7 FR8 GR1 GR2 GR3 GR4 HU1 HU2 HU3 ITC ITF ITG LT0 LU0 LV0 MT0 NO0 PL1 PL2 PL3 PL4 PL5 PL6 R01 R02 R03 R04 SE1 SE2 SE3 SK0 lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl -10 -5 0 5 Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Source: EU-SILC (authors’ calculations). -10 -5 0 5 51 Supplementary material Chapter 1 Figure S5: Excluding individual regions from the sample (ISCED 3–4) -25 -20 -15 -10 -5 0-25 -20 -15 -10 -5 0 AT1 AT2 AT3 BE1 BE2 BE3 BG3 BG4 CY0 CZ0 DK0 EE0 ES1 ES2 ES3 ES4 ES5 ES6 ES7 FR1 FR2 FR3 FR4 FR5 FR6 FR7 FR8 GR1 GR2 GR3 GR4 HU1 HU2 HU3 ITC ITF ITG LT0 LU0 LV0 MT0 NO0 PL1 PL2 PL3 PL4 PL5 PL6 R01 R02 R03 R04 SE1 SE2 SE3 SK0 lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Source: EU-SILC (authors’ calculations). Figure S6: Excluding individual regions from the sample (ISCED 5) AT1 AT2 AT3 BE1 BE2 BE3 BG3 BG4 CY0 CZ0 DK0 EE0 ES1 ES2 ES3 ES4 ES5 ES6 ES7 FR1 FR2 FR3 FR4 FR5 FR6 FR7 FR8 GR1 GR2 GR3 GR4 HU1 HU2 HU3 ITC ITF ITG LT0 LU0 LV0 MT0 NO0 PL1 PL2 PL3 PL4 PL5 PL6 R01 R02 R03 R04 SE1 SE2 SE3 SK0 lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl -20 -10 0 10 20 -20 -10 0 10 20 Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Source: EU-SILC (authors’ calculations). IAB-Bibliothek 367 52 The cohort size-wage relationship in Europe Figure S7: Excluding individual years from the sample ISCED 0–2 ISCED 5ISCED 3–4 Figure S8: Excluding individual experience groups from the sample -20 -10 0 10 20 2004 2005 2006 2007 2008 2009 2010 2004 2005 2006 2007 2008 2009 2010 2004 2005 2006 2007 2008 2009 2010 lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl -10 -5 0 5 -30 -20 -10 0 Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Source: EU-SILC (authors’ calculations). -20 -10 0 10 20 0 1 2 3 4 5 6 7 8 9 10 11 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl lower 95% Cl coefficient upper 95% Cl -10 -5 0 5 -30 -20 -10 0 Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Notes IV2 coefficient from full model (solid) and 95% confidence interval (dashed) Source: EU-SILC (authors’ calculations). ISCED 0–2 ISCED 3–4 ISCED 5 53 Supplementary material Chapter 1 Next, we show the results of a number of changes in the specification of the model. Table S1 shows the results that are obtained when a double-log specification is estimated: cohort size continues to have a significant negative effect for educational group ISCED 3–4 with an elasticity of approximately -0.3. Table S1: Double-log specification ISCED 0–2 ISCED 3–4 ISCED 5 OLS IV1 IV2 OLS IV1 IV2 OLS IV1 IV2 Cohort size (log) -0.03 (0.03) -0.01 (0.08) -0.04 (0.07) -0.01 (0.03) -0.10 (0.09) -0.29** (0.12) -0.07 (0.06) -0.10 (0.23) -0.05 (0.25) N(inds) 7,364 7,364 7,364 19,785 19,785 19,785 4,973 4,973 4,973 N(cells) 2,180 2,180 2,180 2,499 2,499 2,499 1,338 1,338 1,338 N(clusters) 596 596 596 553 553 553 319 319 319 R20.41 0.41 0.41 0.58 0.58 0.58 0.42 0.42 0.42 Control variables are included. ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors are clustered at the region-experience level. N(inds): number of individual observations. N(cells): number of region-experience-year cells. N(clusters) indicates the number of clusters. The use of experience dummies rather than its first two polynomials yields very similar results to those shown in the paper (cf. Table S2). Table S2: Experience dummies ISCED 0–2 ISCED 3–4 ISCED 5 OLS IV1 IV2 OLS IV1 IV2 OLS IV1 IV2 Cohort size -1.46 (1.12) -0.95 (2.97) -2.42 (2.49) -0.30 (0.96) -4.33 (3.47) -12.45*** (4.60) -1.91 (1.54) -3.68 (7.43) -1.91 (7.71) ME (std) -1.75% -1.14% -2.91% -0.25% -3.63% -10.43%*** -2.33% -4.49% -2.33% N(inds) 7,364 7,364 7,364 19,785 19,785 19,785 4,973 4,973 4,973 N(cells) 2,180 2,180 2,180 2,499 2,499 2,499 1,338 1,338 1,338 N(clusters) 596 596 596 553 553 553 319 319 319 R20.42 0.42 0.42 0.58 0.58 0.57 0.42 0.42 0.42 Control variables are included. ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors are clustered at the region-experience level. N(inds): number of individual observations. N(cells): number of region-experience-year cells. N(clusters) indicates the number of clusters. ME(std) shows the percentage change in the hourly wage given a change in cohort size by one standard deviation. In the paper the individual sampling weights are adjusted so as to ensure that the estimated size of a region-year-age cell coincides with the values provided by Eurostat. Table S3 shows that when this is not done and the initial weights are used instead, cohort size retains its negative effect for educational group ISCED 3–4, but the coefficient as well as its marginal effect decrease in magnitude. IAB-Bibliothek 367 54 The cohort size-wage relationship in Europe Table S3: Unadjusted weights ISCED 0–2 ISCED 3–4 ISCED 5 OLS IV1 IV2 OLS IV1 IV2 OLS IV1 IV2 Cohort size -1.43 (1.22) -0.78 (3.98) -2.28 (2.95) -0.39 (0.92) -3.00 (3.30) -8.64** (3.91) -1.07 (1.50) 1.33 (6.91) 2.96 (6.88) ME (std) -1.63% -0.89% -2.60% -0.31% -2.40% -6.88%** -1.23% 1.53% 3.41% N(inds) 7,364 7,364 7,364 19,785 19,785 19,785 4,973 4,973 4,973 N(cells) 2,180 2,180 2,180 2,499 2,499 2,499 1,338 1,338 1,338 N(clusters) 596 596 596 553 553 553 319 319 319 R20.40 0.40 0.40 0.58 0.58 0.57 0.40 0.40 0.40 Control variables are included. ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors are clustered at the region-experience level. N(inds): number of individual observations. N(cells): number of region-experience-year cells. N(clusters) indicates the number of clusters. ME(std) shows the percentage change in the hourly wage given a change in cohort size by one standard deviation. In order to provide a better measure of the supply of individuals in an experience group to the labour market the cohort-size variable that is used in the paper takes into consideration only those individuals that are employed or unemployed. If instead all individuals are included in the construction of the cohort-size variable irrespective of their labour-market status, very similar results are obtained as shown in Table S4. Table S4: Cohort-size variable based on all observations in a region-year-experience cell ISCED 0–2 ISCED 3–4 ISCED 5 OLS IV1 IV2 OLS IV1 IV2 OLS IV1 IV2 Cohort size -0.18 (1.47) -0.95 (3.05) -2.33 (2.45) -1.16 (0.92) -4.41 (3.90) -12.23*** (4.66) -0.50 (1.47) -3.24 (6.62) -1.69 (6.67) ME (std) -0.20% -1.05% -2.56% -1.02% -3.85% -10.70%*** -0.57% -3.67% -1.92% N(inds) 7,364 7,364 7,364 19,785 19,785 19,785 4,973 4,973 4,973 N(cells) 2,180 2,180 2,180 2,499 2,499 2,499 1,338 1,338 1,338 N(clusters) 596 596 596 553 553 553 319 319 319 R20.41 0.41 0.41 0.58 0.58 0.58 0.42 0.42 0.42 Control variables are included. ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors are clustered at the region-experience level. N(inds): number of individual observations. N(cells): number of region-experience-year cells. N(clusters) indicates the number of clusters. ME(std) shows the percentage change in the hourly wage given a change in cohort size by one standard deviation. The specific form of the cohort-size variable has already been used in the literature (Wright, 1991; Brunello, 2010) and reflects the idea that individuals are substitutable with those who are slightly more or less experienced than themselves, while the degree of substitutability is assumed to decrease with the difference in 55 Supplementary material Chapter 1 experience. However, this assumption as well as the inclusion of individuals that differ by at most two years of experience can be argued to be arbitrary. We further assess this issue by specifying alternative cohort-size variables which differ in terms of the number of adjacent experience groups as well as in terms of the use of weights. Specifically, we compute a weighted sum across three experience groups (Equation S1) as well as sums of one, three and five experience groups (Equations S2–S4). In the latter case individuals in the included experience groups are assumed to be perfectly substitutable. CS_2jert = (1/4)Nj – 1, ert + (2/4)Njert + (1/4)Nj + 1, ert Nert [S1] CS_3jert = Njert Nert [S2] CS_4jert = Nj – 1, ert + Njert + Nj + 1 , ert Nert [S3] CS_5jert = Nj – 2, ert + Nj – 1, ert + Njert + Nj + 1, ert + Nj + 2 , ert Nert [S4] Tables S5–S8 contain the corresponding regression results. In each case the coefficient remains negative and significant for ISCED 3–4, but since the means and standard deviations of these variables differ from those of the initial cohort-size variable, the magnitude of the effect can be better compared by looking at the marginal effects rather than at the coefficients. This effect is larger for the three-year weighted average and particularly when only the own experience group is used. The results of the latter specification especially should be treated with caution: the size of region-year-experience-education groups is estimated from EU-SILC data and its accuracy clearly depends on the sampling. In small cells minor changes in the number of observations can have a profound effect on the estimated cohort-size variable. Comparable results to those presented in the paper are found for the three-year sum and a smaller effect for the five-year sum. IAB-Bibliothek 367 62 Regional population structure and young workers’ wages permanently and eventually fell below replacement level. Coupled with increases in life expectancy, these processes are having a substantial effect on the age structure of Germany’s population as evidenced by the ongoing increases in the size of older age groups at the expense of younger ones. Between 1990 and 2010 the ratio of the working-age to the total population fell by over three percentage points, a downward trend that is expected to be exacerbated by the entry into retirement of the large post-World War II birth cohorts. Moreover, demographic change has affected the age composition of the working-age population: while the share of individuals aged 15–24 in the working-age population increased between 2000 and 2010, this development is expected to reverse in the near future with the youth share projected to fall by 2.5 percentage points between 2010 and 2025. The implications of these changes – the combination of a shrinking and ageing population – for the future standard of living constitutes a widely discussed area of research (see Börsch-Supan, 2013). In this context, the question of how labour productivity will be affected by the changes in the population-age structure will be of prime importance (see Bloom and Sousa-Poza, 2013). Likewise, the sustainability of health care and public pension systems in light of demographic pressure has received considerable attention (see Arnds and Bonin, 2002; Jimeno et al., 2008). The objective of this paper is to empirically analyse the impact of changes in the size of the youth population within regional labour markets on the wages of young workers. In the light of the projected population developments, this type of analysis is relevant as it provides a basis for evaluating how demographic processes can be expected to affect the wages of future cohorts of young workers. Given its focus, this paper belongs to a larger body of literature that analyses the effects of changes in the age structure on labour-market outcomes. In addition to wage adjustments, a considerable amount of research has addressed the impact on age-specific (un-) employment (Zimmermann, 1991; Shimer, 2001; Skans, 2005; Biagi and Lucafora, 2008; Ochsen, 2009; Garloff et al., 2013; Moffat and Roth, 2016b) and educational attainment (Connelly, 1986; Stapleton and Young, 1988; Fertig et al., 2009). While wage differences and wage trends between different cohorts in Western Germany are documented in Fitzenberger (1999), his analysis does not focus on the consequences of changes in the age structure, which is the concern of this paper. In a world with a single type of labour input, an increase in the size of the labour force will lead to an outward shift of the labour supply curve. If the labour market works in a way that the wage rate adjusts so as to equate the demand for and the supply of labour and diminishing marginal productivity implies a downward-sloping labour demand curve, the effect of an increase in the labour force will be a lower equilibrium wage rate. If instead labour inputs are not homogenous but rather only 63Chapter 2 Introduction imperfectly substitutable across age groups, the effects of a change in age-specific labour supply will – depending on the degree of substitutability – be concentrated on the members of that age group. Within such a framework, an increase in the share of young individuals should be accompanied by a decrease in their wages. Our contribution is threefold. First, our assessment of the relationship between the youth share and young workers’ wages in Western Germany addresses the lack of recent empirical evidence on this topic. Second, we use functional entities in order to identify the size of the youth population within an actual labour market rather than within an administrative unit as is done by earlier studies, which reduces the potential for measurement error in this variable. Third, we assess the channels through which changes in the size of the youth population affect young workers’ wages by controlling for industrial and occupational upor downgrading. Gertler and Trigari (2009) argue that individuals have a better chance of moving into higher-paying industries, firms or jobs during boom periods than during recessions. We propose that a similar argument can be made with respect to agegroup size, as increased competition may lead individuals to take up positions in lower-paying industries or occupations than they would have done, had they been part of a smaller age group. In order to distinguish between the direct and the selection-related, indirect effect of belonging to a larger age group, we compare the estimated wage effect of the youth share from models that exclude or include detailed information about an individual’s industrial and occupational affiliation. In our model the effect that the regional youth share has on the wages of young workers is identified solely through the within-variation of this variable. However, as the relative size of the youth population within a labour market is potentially endogenous due to migration into high-wage areas, an instrumental variables (IV) identification strategy is employed: within a given region the instrument is defined as the share of individuals that are fifteen years younger and that are observed fifteen years earlier than the age group of the endogenous regressor. We find that the youth share has a statistically significant negative effect on the wages of young workers. Specifically, an increase by one percentage point is predicted to decrease wages by 3% in our baseline model. When using a district-based measure of the youth-share variable, the estimated coefficients are smaller by between 13% and 48%. Finally, we find that controlling for an individual’s industry and, particularly, occupation reduces the estimated wage decrease from 3% to 2%, which suggests that a substantial part of the negative effect of age-group size is the result of individuals in larger age groups being more likely to be employed in lower-paying occupations. According to these results, future generations of young workers can expect to benefit from demographic developments. Specifically, a decrease in the youth share by 2.5 percentage points, IAB-Bibliothek 367 64 Regional population structure and young workers’ wages as projected to occur between 2010 and 2025, would be predicted to lead to an increase in young workers’ wages of about 5%, ceteris paribus. The remainder of the paper is structured as follows. Section 2 addresses the relationship between age structure and wage outcomes and reviews the relevant theoretical and empirical literature. Section 3 provides descriptive statistics on the youth population in Germany. The empirical analysis is the topic of Section 4, while Section 5 discusses the regression results. Section 6 presents the conclusion. 2 Population structure and wages Differently aged workers are not perfectly substitutable. Age can be expected to be correlated with a worker’s set of skills, which in turn affects his suitability for different tasks. First, age is a good predictor for work experience, and, ceteris paribus, more experienced workers will usually have more firm-specific, occupationspecific, industry-specific or general human capital. If this type of knowledge is relevant for on-the-job performance, differently aged workers can be expected to be only imperfectly substitutable. Indeed, Welch’s (1979) career-phase model can be interpreted as an example of a model in which imperfect substitutability arises from differences in firm-specific human capital. Second, jobs vary with respect to the tasks that they contain and therefore also concerning the abilities that workers are required to have in order to perform these tasks. Older workers may be less easily substitutable for younger workers in occupations requiring physical or certain types of cognitive skills (Mazzonna and Perracchi, 2012). As a consequence of imperfect substitutability a change in the relative size of an age group will mainly affect the labour market outcomes of the members of that group. As a starting point to analysing the effects of a change in the size of a specific age group on the wages of its members, it is useful to assume a production function with differently aged workers as distinct factors of production (see Card and Lemieux, 2001; Fitzenberger and Kohn, 2006). In the benchmark case of a perfectly competitive labour market, in which each factor of production is paid the monetary value of his marginal product, a change in the supply of a specific production factor will cause the wage to adjust in a way that the market is again cleared. In the case of each factor of production exhibiting diminishing marginal productivity, an increase in the size of an age group will reduce the wages paid to its members. Labour markets, however, do not necessarily clear. The existence of minimum or efficiency wages as well as collective wage bargaining are possible sources that can prevent the wage rate from fully adjusting in response to a change in labour supply, while the coexistence of unemployment and vacancies provides evidence against the existence of a market-clearing equilibrium as predicted by the benchmark model 65Chapter 2 Population structure and wages of a competitive labour market. Existing theoretical models, however, suggest that even in the absence of clearing labour markets, changes in the relative supply of an age group will have an effect on age-specific wages (Michaelis and Debus, 2011). The extant empirical literature, though differing with respect to the time periods and countries (or regions) under study, the model specification and identification strategy, provides evidence that increases in the size of an age group are associated with depressed wage outcomes for the members of that group.2 Early studies using US data estimate a negative relationship between the relative size of an age group and the average wages that are earned by individuals within that group for different levels of educational qualification (Welch, 1979; Berger, 1985). Alternatively, Freeman (1979) finds a negative effect of the young-to-old population ratio on the average wages of young workers relative to those of old workers. The existence of a negative effect of age-group size is also supported by evidence from Sapozhnikov and Triest (2007). Most recently, Morin (2015) exploits an exogenous shock to the supply of high-school graduates in Canada due to a reform of the secondary schooling system and finds negative cohort-size effects on wages. Empirical evidence from Europe is scarcer but also supports the hypothesis that wages earned in larger age groups are depressed compared to those of smaller age groups (see Wright, 1991, for the UK and Brunello, 2010, as well as Moffat and Roth, 2016a, for a sample of European countries). A drawback with respect to identifying the effect of interest is that the size of an age group within a given spatial unit is arguably endogenous due to selfselection of individuals into high-wage areas. Korenman and Neumark (2000) proposed birth rates as an instrument, while other authors have since used the lagged relative size of age groups as exogenous predictors (Skans, 2005; Garloff et al., 2013; Moffat and Roth, 2016a and 2016b). Whereas cross-country migration might be deemed too small to influence the size of nationally defined age groups, endogeneity resulting from self-selection through migration becomes a larger concern when the spatial units that are used to construct the measure of population structure are defined at a sub-national level. While many empirical studies in this field of research have used measures of population structure at the national level, it appears questionable whether a country indeed constitutes the appropriate delineation of a labour market. If individuals are restricted in their mobility or if awareness of job openings in other regions decreases with distance, a nationally defined youth-share variable groups together young individuals that are not active in the same labour market and that are hence not 2 Notable exceptions can be found in the migration literature where many studies conclude that natives’ wages are not negatively affected by age-specific immigration (Ottaviano and Peri, 2012). A possible explanation for this finding is that migrants are complements rather than substitutes for native labour. IAB-Bibliothek 367 66 Regional population structure and young workers’ wages substitutable for one another. Such a variable would be subject to measurement error if labour markets existed at a sub-national level and the size of the youth population varied across them. And while more recent studies have made use of administrative units at a sub-national level, so-constructed youth-share variables may still be measured with error as administrative units are generally not delineated in a way as to coincide with actual labour markets, meaning that they would not necessarily capture the relative supply of young labour that is relevant for the determination of a young worker’s wages. To address this issue, we employ the functional labour-market regions that are defined by Eckey et al. (2006). These regions consist of one or more districts (Kreise) and are constructed on the basis of observed commuting flows with a typical labour-market region combining an economic centre with the surrounding Umland from which people commute to work in the centre. They approximate selfcontained local labour markets in as far as they aim to maximise the overlap between the population living and working within such a region. Functional units therefore provide a better measure of the size of the youth population in an actual labour market than administrative units. The self-contained nature of these units also reduces the need to consider the youth population in surrounding labour markets as a factor determining the wages of young workers in a given region. It should be noted that changes in the age structure of the population do not necessarily imply changes in age-specific labour supply as participation rates as well as the number of hours worked could in principle adjust in a way as to completely counteract changes in age-group size. However, such a reaction seems unlikely as empirical evidence suggests that male labour supply is inelastic – at the extensive and the intensive margin – to changes in the wage rate (Blundell and MaCurdy, 1999). More specifically, Garloff et al. (2013) show that a counteracting development in participation rates has not taken place in Germany in response to changes in the age structure at the national level in recent years. 3 Youth-population structure in Western Germany This section provides information about the development of the working-age (15–64) and the youth population (15–24) in Western Germany at the national level and at the level of the labour-market region. Figure 1 shows the absolute size of both populations at five-year intervals between 1995 and 2040. While the actual values are shown up to the year 20103, subsequent developments represent projections based on the variant Untergrenze der mittleren Bevölkerung, which 3 Data comes from the Federal Statistical Office and has been obtained through the following link: https://www. genesis.destatis.de/genesis/online/link/tabellen/12411* 67Chapter 2 Youth-population structure in Western Germany assumes an annual net immigration of 100,000 individuals and a fertility rate of 1.4 and which represents the lower bound of corridor within which population development is expected to take place (Statistisches Bundesamt, 2010).4 Except for a small increase between 1995 and 2000, the working-age population has been shrinking steadily and is projected to continue decreasing in size over the coming decades. By 2040 it will have fallen by almost 25% compared to its 2010 value, which reflects the effect of the large post-World War II birth cohorts reaching retirement age. In contrast, the number of young individuals grew by half a million between the years 2000 and 20105, but this development is expected to reverse in the near future with the size of the age group 15–24 projected to fall continuously until 2040. Reflecting changes in these two populations’ relative rate of growth, the youth share, i.e. the size of the population aged 15–24 relative to the working-age population, displays a cyclical development: from 2000 to 2010 the share of young individuals expanded by approximately one percentage point (equivalently, 7%). However, as the youth population is expected to decrease at a faster rate than the working-age population, its share is projected to fall by 4 The upper bound of this corridor ( Obergrenze der mittleren Bevölkerung ) differs by assuming that annual net immigration will increase steadily to 200,000 in the year 2020 before plateauing at that level. Despite this difference the projection for the youth share is very similar (the largest difference between both projections amounts to 0.25 percentage points in the year 2040). 5 These age groups are the children of the large post-World War II birth cohorts. This increase therefore reflects the large size of the parental generation. Figure 1: Development of the population and the youth share at the national level Source: Federal Statistical Office. 0.15 0.16 0.17 0.180.14 Youth share 0 10,000 20,000 30,000 40,000 Population (in 1,000s) 1995 2000 2005 2010 2015 2020 2025 2030 2035 2040 Population (15–64) Population (15–24) Youth share IAB-Bibliothek 367 68 Regional population structure and young workers’ wages 2.5 percentage points (equivalently, 15%) between 2010 and 2025. At the national level, the increase in the youth share during most of the sample period therefore contrasts with its projected development in the immediate future, which implies that changing demographics may contribute positively towards the development of young workers’ wages in the coming years. Figure 2 illustrates the existing regional heterogeneity in the share of individuals aged between 15 and 24 in the working-age population by reporting the value of this variable for the West-German regional labour markets. The extent of cross-sectional variation in the youth-share variable is revealed for the year 1995 in the top left map, in which the labour-market regions are grouped into quartiles based on the size of the youth share. Compared to a value of about 16% at the national level, the regional youth share varies between 14% and 21%. Figure 2: Variation in the youth share (15–24) at the regional level Source: Federal Statistical Office (population data) and Federal Institute for Research on Building, Urban Affairs and Spatial Development (geodata). 1995 2000 2005 2010 Share 15–24 (0.177,0.204) (0.170,0.177) (0.160,0.170) (0.145,0.160) 69Chapter 2 Empirical analysis The other maps show the cross-sectional variation in the youth-share variable for the years 2000, 2005 and 2010, respectively. Moreover, they reveal the within-region variation in this variable, i.e. its development over time (to allow for a comparison of the different years, the same intervals are chosen as for the year 1995). Reflecting the drop in the national youth share in the year 2000, the share has also generally fallen at the regional level as illustrated by a number of regions that were in the fourth or third quartile in 1995 now being in the third or second quartile, respectively. Likewise, an increasing number of regions are registered in higher quartiles in the years 2005 and 2010, reflecting the increase in the youth share at the national level. 4 Empirical analysis The different steps of empirically analysing the relationship between the youth share and young workers’ wages are the subject of this section: the relevant datasets are introduced in the first part, which is followed by a description of how the sample is constructed and how the model’s main variables are defined. The final part discusses the empirical model and the identification strategy. 4.1 Data Three data sources are used for the empirical analysis. The first source is population data for Germany on the regional level according to age groups which is used to construct the relative size of the youth population within a regional labour market. The information reported by the statistical offices refers to the end of the year (31 December). There is no information beyond age and sex in these data. Particularly, there is no information available on the educational composition. Corrections have to be made to account for changes in the delineations of municipalities and districts which results in a dataset that is spatially consistent over time back until 1978. However, the available age brackets differ for the time before and after 1985. Second, we use statistics from the Federal Employment Agency (FEA) to gather information on employment numbers and rates as well as unemployment rates. Employment numbers and rates can be obtained at the level of the labour-market regions starting in 1987 for employment at place of work and from 1999 for employment at place of residence. The data is available by single-age cohorts, sex and education and refers to the middle of the year (30 June), since those values are typically close to yearly averages. In order to better compare the results from a model using a youth-share variable based on an individual’s place of residence with those derived from an individual’s place of employment, the year 1999 is chosen as the start of the sample period. IAB-Bibliothek 367 70 Regional population structure and young workers’ wages The final source is the Stichprobe der Integrierten Erwerbsbiografien (SIAB), a large micro-dataset from the Institute for Employment Research (IAB), that includes information on a 2% random sample of all individuals in Germany that were employed, unemployed or participating in measures of active labour-market policy between 1975 and 2010 (civil servants and the self-employed are excluded). For employed individuals in the dataset we have information on their employment relationship on a daily basis. Moreover, it contains a wealth of additional information that we use in part as control variables. The data further contains information about an employee’s place of residence and place of employment, though the former only becomes available in 1999. A detailed description of the dataset can be found in vom Berge et al. (2013). 4.2 Sample and descriptive statistics The observations contained in SIAB refer to spells of an individual (e.g. an employment spell) with given start and end dates as well as characteristics of the spell (e.g. the average daily wage earned during this period). We use the settingup routines by Eberle et al. (2013) to transform the structure of the data so that it contains data from a single spell per individual and year. In doing so, we choose 15 June as the annual reference date, which means that only those spells are retained that include the reference date in a given year. As employers are required to report the wages of their employees once a year and this is typically done on 31 December, the longest spells run from 1 January to 31 December in a given year. Using 15 June as the reference date implies that spells starting and ending before (or after) 15 June within a given year are not being considered. This specific reference date is chosen because June values of employment figures are usually close to annual averages, while the middle of the month is used to avoid any endof-calendar-month effects. However, the results are robust to using 31 December as the reference date.6 The sample covers the period 1999–2010 and consists of regularly employed (sozialversicherungspflichtig Beschäftigte) males who are between 15 and 24 years old. Individuals in vocational training are excluded because the mechanisms determining their remuneration are considered to be different from the rest of the labour market. As there is no information about the number of hours worked in the data, the sample is further restricted to full-time employees. While 95% of the observations have one full-time job, some observations hold other jobs in addition to being full-time employed, e.g. 3% of observations are also in minor 6 The results of this and all other robustness checks can be found in the Supplementary Material to this paper. 71Chapter 2 Empirical analysis employment (geringfügige Beschäftigung). In such a case only information about the first full-time job is retained.7 We do not restrict employment spells to have a minimum duration. However, the results are robust to keeping only observations with employment spells of at least 90 days in the sample. The model’s dependent variable is an individual’s inflation-adjusted daily wage including social security contributions and taxes.8 The reported wage is censored at the value of the corresponding year’s upper social security threshold; but given that our sample is restricted to individuals aged between 15 and 24 only a small fraction of observations will have wages above the threshold, and since imputation procedures (see Gartner, 2005) suggest that in such a case the true wage values are close to the censoring value, we use the censored wage for these observations. At the other end of the spectrum, we also observe unrealistically low daily wages. To remove these observations we truncate the wage distribution at twice the value of the minor-employment threshold (Geringfügigkeitsgrenze) – an approach that has also been taken by other authors working with the same data source (e.g. Gürtzgen, 2016). This implies that observations with wages of less than 650 Euro per month (21.26 Euro per day or, alternatively, 2.57 Euro per hour, assuming an eight-hour working day) between 1999 and 2002 or less than 800 Euro per month (26.28 Euro per day or 3.29 Euro per hour) between 2003 and 2010 are dropped.9 The main variable of interest is the youth share, which measures the number of individuals aged between 15 and 24 relative to the number of working-age individuals (ages 15–64) within a regional labour market as defined by Eckey et al. (2006).10 Due to limitations pertaining to the availability of population data preceding re-unification, our empirical analysis is restricted to the 108 labourmarket regions (313 districts) of Western Germany. This restriction is unfortunate: the demographic processes that have seen the youth share in Eastern Germany fall 7 For individuals holding more than one job at the same time it would in principle be possible to use total earnings from all jobs rather than just the wage earned in one job as the relevant dependent variable. We abstain from doing so as our focus is on how the supply of young workers affects the wages earned in a particular job. Similar results to those shown in Table 1 are obtained when observations with more than a full-time job are removed from the sample. 8 Inflation-adjustment is done using the consumer price index (base year: 2010). The data comes from the Federal Statistical Office and has been obtained through the following link: https://www.destatis.de/DE/ZahlenFakten/ GesamtwirtschaftUmwelt/Preise/Verbraucherpreisindizes/Tabellen_/VerbraucherpreiseKategorien.html?cms_ gtp=145110_slot%253D2&https=1 9 If individuals with wages below the specified thresholds are not excluded from the analysis, the youth-share coefficients are smaller in size and less significant. Compared to the sample used in the empirical analysis of this paper, individuals below the threshold are more likely to have a lower secondary education without apprenticeship training (56% compared to 19%) and are employed in firms with on average a smaller number of employees (446 compared to 970), whereas the average size of the youth share is similar. In addition to measurement error in the wages, the decrease in the effect of the youth share might also be due to the wages of this group being less responsive to changes in the supply of young workers, possibly because they are downward-rigid due to institutional constraints (e.g. sector-specific minimum wages). 10 Similar results are obtained when we use an employment-based youth-share variable that is defined as the number of employed youths aged 15–24 relative to the workforce. IAB-Bibliothek 367 78 Regional population structure and young workers’ wages earnings, which is in line with evidence by Lehmer and Möller (2010). Finally, the effects of the unemployment rate and the share of young unemployed individuals are small. The youth-unemployment variable draws a negative coefficient in the 2SLS estimations, but in contrast to findings by Baltagi and Blien (1998) its effect is not statistically significant. Related studies have used administrative units at the sub-national level as the basis for constructing population variables. As discussed in Section 2, the drawback of such an approach is that these units do not necessarily represent actual labour markets and that, consequently, the size of age groups within a given labour market is potentially measured with error (see the Supplementary Material for a discussion). We assess the effect of using administrative rather than functional units by estimating Equation 1 on the basis of a district-specific youthshare variable. The results are shown in Table 2. Table 2: District-based youth share variable Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.31 (0.45)*** -2.79 (0.81)*** -0.12 (0.42) -1.50 (0.92) Dummies Year District Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage regression Instrument First-stage test statistics F-statistic Shea’s partial R2 – – – 0.44 (0.00)*** 300.92*** 0.27 – – – 0.43 (0.00)*** 181.95*** 0.22 Observations Individuals District-year cells Districts (clusters) 107,351 3,756 313 107,351 3,756 313 107,351 3,756 313 107,351 3,756 313 R20.25 0.25 0.26 0.26 ME(stdev) -1.80%*** -3.84%*** -0.17% -2.18% Cluster-robust standard errors in parentheses (clustered at the district level). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. The 2SLS coefficients of the youth-share variables remain negative and larger in absolute value than the corresponding OLS estimates, but only the specification referring to an individual’s place of residence produces statistically significant results. However, compared to the results of Table 1, using districts rather than labour-market regions leads to an underestimation of the youth share’s negative 79Chapter 2 Results effect: the point estimates referring to the place of residence are smaller by 13%, while the size of the coefficient for the place of employment drops by almost 50%.17 The increased discrepancy between the youth-share coefficients of these specifications reflects the fact that individuals are more likely to live and work in different districts than is the case for labour-market regions. It can be shown that the negative and significant youth-share coefficients of Table 2 are driven by those districts in which individuals in the sample are more likely to live than to work. At the same time, the average absolute difference between the district-based youth-share variable and its value at the corresponding labourmarket region is smaller for these districts, which suggests that measurement error in the size of the youth-share variable is less pronounced. The fact that these districts account for a larger fraction of observations in the place-of-residence specification suggests that size and significance of the youth-share coefficients will be less affected in that specification. It turns out that individuals in the sample are more likely to work in cities and to live in rural areas. The rationale behind the above argument could therefore be that the youth share within a city-district provides only an inaccurate measure of the size of the youth population that is relevant for the determination of wage outcomes as cities will also draw workers from surrounding districts. As discussed in Section 1, the size of an individual’s age group could have an effect on the conditions of his employment. Specifically, if young workers in larger age groups are more likely to be in positions in lower-paying occupations or industries, the estimated wage effect of the youth share in Table 1 would be confounded by these types of selection effects. In particular, the negative effect would be overestimated. To address this issue, we successively add indicator variables to the model of Equation 1 which are derived from two-digit codes referring to an individual’s industry and occupation. The results are shown in Table 3 for the placeof-residence specification and in Table 4 for the place of employment. In both cases we find that adding industry and, especially, occupation indicators has a sizeable impact on the estimated youth-share effects compared to the baseline specification: the inclusion of industry indicators decreases the size of the 2SLS coefficients by 12% (place of residence) and 4% (place of employment) compared to the results of the baseline model, while the reduction resulting from adding occupation indicators is considerably larger at 40% and 30%, respectively. Similar results are obtained when both sets of indicator variables are used. Moreover, it can be seen that when dummies for industry or occupation are added the difference in 17 Due to the higher variance of the district-based youth-share variable the proportional changes in the marginal effects for a change of one standard deviation are less pronounced. IAB-Bibliothek 367 80 Regional population structure and young workers’ wages the size of the 2SLS coefficients between the place of residence and the place of employment decrease in magnitude. This supports the argument that once labourmarket regions are used as the spatial entities from which the youth-share variable is constructed both types of places produce similar results. Table 3: Industry and occupation indicators (place of residence) Dependent variable: log real daily earnings Baseline +industry +occupation +industry +occupation Youth share (2SLS) Youth share (OLS) -3.22 (0.97)*** -1.46 (0.63)* -2.81 (0.92)*** -1.36 (0.57)* -1.88 (0.91)* -0.90 (0.60) -1.89 (0.86)* -0.97 (0.53)† Dummies Year Labour-market region Industry Occupation Control variables Yes Yes No No Yes Yes Yes Yes No Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes First-stage regression Instrument First-stage test statistics F-statistic Shea’s partial R2 0.46 (0.00)*** 136.60*** 0.32 0.46 (0.00)*** 137.17*** 0.32 0.46 (0.00)*** 137.32*** 0.32 0.46 (0.00)*** 137.70*** 0.32 Observations Individuals Labour-market region-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R2 (2SLS) R2 (OLS) 0.24 0.24 0.46 0.46 0.40 0.40 0.51 0.51 ME(stdev, 2SLS) ME(stdev, OLS) -4.05%*** -1.84%* -3.54%*** -1.71%* -2.36%* -1.14% -2.39%* -1.22%† Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. While in the baseline model an increase in the size of the youth-share variable by one percentage point was predicted to decrease an individual’s wages by about 3%, ceteris paribus, the size of this effect is reduced once an individual’s industrial and, in particular, occupational affiliation are controlled for. This finding suggests that the estimated youth-share coefficients of the baseline specification were indeed confounded by the positive association between young workers being in larger age groups and being employed in lower-paying industries and occupations. We conclude that in addition to the direct negative effect of the size of the youth share, there is an indirect effect driven by selection into specific industries and occupations. A possible explanation for this finding is that, ceteris paribus, a larger supply of young individuals increases competition for higher-quality jobs, forcing some individuals to take up employment in lower-paying occupations. 81Chapter 2 Results Table 4: Industry and occupation indicators (place of employment) Dependent variable: log real daily earnings Baseline +industry +occupation +industry +occupation Youth share (2SLS) Youth share (OLS) -2.89 (1.22)* -0.85 (0.73) -2.77 (1.06)** -0.89 (0.62) -1.96 (1.08)† -0.52 (0.66) -2.11 (1.01)* -0.62 (0.57) Dummies Year Labour-market region Industry Occupation Control variables Yes Yes No No Yes Yes Yes Yes No Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes First-stage regression Instrument First-stage test statistics F-statistic Shea’s partial R2 0.46 (0.00)*** 131.80*** 0.32 0.46 (0.00)*** 132.31*** 0.32 0.46 (0.00)*** 132.33*** 0.32 0.46 (0.00)*** 132.67*** 0.32 Observations Individuals Labour-market region-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R2 (2SLS) R2 (OLS) 0.24 0.24 0.46 0.46 0.41 0.41 0.51 0.51 ME(stdev, 2SLS) ME(stdev, OLS) -3.67%* -1.09% -3.52%** -1.13% -2.49%† -0.66% -2.68%* -0.79% Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. This interpretation is in line with recent results pertaining to the wage effects of labour market conditions. Kahn (2010) and Brunner and Kuhn (2014) find that adverse labour market conditions (measured by the unemployment rate at the time of labour-market entry) depress wages and increase the probability of employment in lower-quality occupations. Morin (2015) studies the wage effects of the increase in labour supply due to the double cohort of high-school graduates in Ontario and provides evidence that part of the negative wage effect is due to selection into lower-paying occupations. Alternatively, higher-quality jobs may require a specific type of qualification. If the supply of training positions does not adjust to the supply of young individuals, the number of individuals barred from entering higher-paying occupations will increase in larger age groups. The effect of age-group size on selection into industries and occupations certainly warrants further research. The Supplementary Material contains the corresponding output tables for the case in which districts provide the basis for the construction of the youth-share variable. These show that region-specific and district-specific variables continue to produce different results once industry and occupation dummies have been added and also illustrate that the difference between the results of the place-of-residence and the place-of-employment specifications are more pronounced at the district level. IAB-Bibliothek 367 82 Regional population structure and young workers’ wages To assess to what extent the results of Table 1 merely reflect unobserved heterogeneity at the federal state-year level, we add dummy variables for the interaction between federal states and years to the model of Equation 1. Doing so allows us to control for annual shocks that affect states differently and that are relevant for the determination of individual wages, e.g. the effects of macroeconomic shocks may vary between states due to differences in industrial structure. The results displayed in Table 5 suggest that, at least for the place-ofresidence specification, the estimated effects of the youth-share variable in the baseline specification are not driven by unobserved heterogeneity at the state-year level. Table 5: State-by-year interactions Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.01 (0.63) -2.58 (1.27)* -0.46 (0.74) -2.35 (1.56) Dummies Year Labour-market region Federal state-by-year Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage regression Instrument First-stage test statistics F-statistic Shea’s partial R2 – – – 0.43 (0.00)*** 85.39*** 0.26 – – – 0.43 (0.00)*** 84.74*** 0.26 Observations Individuals Labour-market region-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -1.28% -3.25%* -0.58% -2.99% Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. The 2SLS point estimates fall by approximately 20% as part of the explained variation in the earnings variable is now picked up by the additional dummies. The standard errors increase presumably because parts of the variation in the youth-share variable are now explained by the additional dummy variables, which results in less precise estimates. For a similar reason there is a drop in the values of the first-stage F-statistic and the partial R2 of the excluded instrument: the explanatory power of the instrument is reduced as a consequence of including the state-by-year dummies in the first-stage equation. 83Chapter 2 Conclusion 6 Conclusion This paper empirically analyses how changes in the size of the youth population affect the wages of young workers. Under the assumption that differently aged individuals are only imperfectly substitutable because of differences in firmspecific, occupation-specific, industry-specific or general human capital, economic theory predicts that an increase in the size of an age group reduces the earnings of the members of that group. This hypothesis is tested using a sample of young male employees from Western Germany. The demographic forces that are currently changing the age-structure of the German population illustrate the relevance of this analysis. Specifically, the share of young individuals is projected to fall by 2.5 percentage points (equivalently, by 15%) at the national level over the period 2010–2025 following a period of an increasing youth share. Besides providing an analysis of this relationship using recent data from administrative records, this paper makes two additional contributions. First, functional labour-market regions rather than administrative units are used as the spatial entities within which the size of the youth population is measured. These units provide a better measure of the number of young individuals in an actual labour market than administrative units, which are usually not delineated according to economic criteria, and hence of the supply of young labour that is relevant for the determination of young workers’ wages. Use of a youth-share variable based on labour-market regions therefore reduces the potential for measurement error in this variable. Second, we address the channels through which an increase in the supply of young individuals affects their wages by controlling for industrial and occupational upgrading, i.e. for the possibility that changes in the size of the youth population affect the chances of finding employment in higher-paying industries or occupations. The empirical analysis employs an IV approach in order to account for the possibility that the youth share is endogenous due to young individuals migrating into high-wage areas. In line with the hypothesis that increases in age-group size reduce the wages of the members of that group, the 2SLS coefficients are negative and significant: an increase in the youth share by one percentage point is predicted to decrease young workers’ wages by 3%. Consistent with the argument that migration into high-wage regions induces endogeneity, the corresponding OLS estimates are less negative. Estimating our model using a youth-share variable that is based on districts rather than labour-market regions reduces the size of the 2SLS coefficients by either 13% (place of residence) or 48% (place of employment), which is consistent with the hypothesis that the use of administrative units induces measurement error in the youth-share variable. Finally, adding indicators for an individual’s occupation and industry reduces the size of the youth-share IAB-Bibliothek 367 84 Regional population structure and young workers’ wages coefficients from -3% to -2%. We interpret this result as providing evidence for the hypothesis that belonging to a larger age group increases the likelihood of being employed in lower-paying occupations or industries. What are the implications of these findings for the wages of the coming cohorts of young workers in light of Western Germany’s changing demographics? As the youth share is projected to decrease over the coming years, demographic processes appear to be favourable to the development of the wages that young workers can expect in the future. But as the development of population structures is likely to differ between regions, regional variation in the extent to which young workers stand to benefit is to be expected. Finally, it should be borne in mind that these results come from a specific sample consisting of young, male, fulltime employees with a few years of work experience and, predominantly, lower secondary education. Whether the relationship between the youth share and young workers’ wages is similar for other groups, such as females or the highly educated, remains a topic for future research. Acknowledgements The authors would like to thank Stefan Fuchs, Bernd Hayo, John Moffat and Norbert Schanne for their advice and are grateful for comments from the participants of IAB’s regional research network meeting in Aalen, the 54th Conference of the European Regional Science Association (ERSA), the 5th ifo Workshop Arbeitsmarkt und Sozialpolitik, the joint IAB Regional Science Academy workshop in Amsterdam and the 28th Conference of the European Association of Labour Economists (EALE). The Federal Institute for Research on Building, Urban Affairs and Spatial Development kindly provided the shapefile of the German labour-market regions. Our thanks also go to Annie Roth for proofreading an earlier version of this paper. References Angrist, J. and Pischke, J.-S. (2009) Mostly harmless econometrics: an empiricist’s companian, Princeton University Press, Princeton. Arnds, P. and Bonin, H. (2002) Arbeitsmarkteffekte und finanzpolitische Folgen der demographischen Alterung in Deutschland, IZA Discussion Paper No. 667, Institute of Labor Economics, Bonn. Baltagi, B.H. and Blien, U. (1998) The German wage curve: evidence from the IAB employment sample, Economics Letters, 61, 135–42. 85Chapter 2 References Berger, M.C. (1985) The Effect of Cohort Size on Earnings Growth: A Reexamination of the Evidence, Journal of Political Economy, 93, 561–73. Biagi, F. and Lucifora, C. (2008) Demographic and education effects on unemployment in Europe, Labour Economics, 15, 1076–101. Bloom, D.E. and Sousa-Poza, A. (2013) Aging and productivity: Introduction, Labour Economics, 22, 1–4. Blundell, R. and MaCurdy, T. (1999) Labour Supply: A Review of Alternative Approaches, in O. Ashenfelter and D. Card (eds.) Handbook of Labor Economics, Vol. 3., North-Holland, Amsterdam. Börsch-Supan, A. (2013) Myths, Scientific Evidence and Economic Policy in an Aging World, Journal of the Economics of Ageing, 1–2, 3–15. Brunello, G. (2010) The effects of cohort size on European earnings, Journal of Population Economics, 23, 273–90. Brunner, B. and Kuhn, A. (2014) The impact of labor market entry conditions on initial job assignments and wages, Journal of Population Economics, 27, 705–38. Card, D. and Lemieux, T. (2001) Can Falling Supply Explain the Rising Return to College for Younger Men? A Cohort-Based Analysis, Quarterly Journal of Economics, 116, 705–46. Connelly, R. (1986) A Framework for Analyzing the Impact of Cohort Size on Education and Labor Earning, Journal of Human Resources, 21, 543–62. Eberle, J., Schmucker, A. and Seth, S. (2013) Programmierbeispiele zur Datenaufbereitung der Stichprobe der Integrierten Arbeitsmarktbiografien (SIAB) in Stata. Generierung von Querschnittsdaten und biografischen Daten, FDZ Methodenreport No. 04/2013, Institute for Employment Research, Nuremberg. Eckey, H.-F., Kosfeld, R. and Türck, M. (2006) Abgrenzung deutscher Arbeitsmarktregionen, Raumforschung und Raumordnung, 64, 299–309. Fertig, M., Schmidt, C. and Sinning, M. (2009) The Impact of Demographic Change on Human Capital Accumulation, Labour Economics, 16, 659–68. Fitzenberger, B. (1999) Wages and employment across skill groups: An analysis for West Germany, Physica-Verlag: Heidelberg. Fitzenberger, B. and Kohn, K. (2006) Skill Wage Premia, Employment, and Cohort Effects: Are Workers in Germany All of the Same Type?, IZA Discussion Paper No. 2185, Institute of Labor Economics, Bonn. Freeman, R.B. (1979) The effect of demographic factors on age-earnings profiles, Journal of Human Resources, 14, 289–318. Fuchs, M. and Weyh, A. (2014) Demography and unemployment in East Germany. How close are the ties?, IAB Discussion Paper No. 26/2014, Institute for Employment Research, Nuremberg. IAB-Bibliothek 367 86 Regional population structure and young workers’ wages Garloff, A., Pohl, C. and Schanne, N. (2013) Do small labor market entry cohorts reduce unemployment?, Demographic Research, 29, 379–406. Gartner, H. (2005) The imputation of wages above the contribution limit with the German IAB employment sample, FDZ Methodenreport No. 02/2005, Institute for Employment Research, Nuremberg. Gertler, M. and Trigari, A. (2009) Unemployment Fluctuations with Staggered Nash Wage Bargaining, Journal of Political Economy, 117, 38–86. Gürtzgen, N. (2016) Estimating the Wage Premium of Collective Wage Contracts - Evidence from Longitudinal Linked Employer-Employee Data, Industrial Relations, 55, 294–322. Jimeno, J.F., Rojas, J.A. and Puente, S. (2008) Modelling the impact of aging on social security expenditures, Economic Modelling, 25, 201–24. Kahn, L.B. (2010) The long-term labor market consequences of graduating from college in a bad economy, Labour Economics, 17, 303–16. Korenman, S. and Neumark, D. (2000) Cohort Crowding and Youth Labor Markets: A Cross-National Analysis, in D.G. Blanchflower and R.B. Freeman (eds.) Youth Employment and Joblessness in Advanced Countries, University of Chicago Press, Chicago. Lehmer, F. and Möller, J. (2010) Interrelations between the urban wage premium and firm-size wage differentials: a microdata cohort analysis for Germany, The Annals of Regional Science, 45, 31–53. Mazzonna, F. and Peracchi, F. (2012) Ageing, cognitive abilities and retirement, European Economic Review, 56, 691–710. Michaelis, J. and Debus, M. (2011) Wage and (un-)employment effects of an ageing workforce, Journal of Population Economics, 24, 1493–511. Mincer, J. (1958) Investment in Human Capital and the Personal Income Distribution, Journal of Political Economy, 66, 281–302. Moffat, J. and Roth, D. (2016a) The Cohort Size-Wage Relationship in Europe, Labour, forthcoming. Moffat, J. and Roth, D. (2016b) Cohort size and youth labour-market outcomes: the role of measurement error, IAB Discussion Paper No. 37/2016, Institute for Employment Research, Nuremberg. Morin, L.-P. (2015) Cohort size and youth earnings: Evidence from a quasiexperiment, Labour Economics, 32, 99–111. Moulton, B.R. (1990) An Illustration of a Pitfall in Estimating the Effects of Aggregate Variables on Micro Units, Review of Economics and Statistics, 72, 334–38. Ochsen, C. (2009) Regional labor markets and aging in Germany, Thünen-Series of Applied Economic Theory Working Paper No. 102, Department of Economics, Faculty of Economics and Social Sciences, University of Rostock, Rostock. 87Chapter 2 References Ottaviano, G. and Peri, G. (2012) Rethinking the effect of immigration on wages, Journal of the European Economic Association, 10, 152–97. Polachek, S.W. (2008) Earnings Over the Life Cycle: The Mincer Earnings Function and Its Applications, Foundations and Trends in Microeconomics, 4, 165–272. Sapozhnikov, M. and Triest, R.K. (2007) Population Aging, Labor Demand, and the Structure of Wages, Research Department Working Paper 07-8, Federal Reserve Bank of Boston, Boston. Shimer, R. (2001) The Impact of Young Workers on the Aggregate Labor Market, The Quarterly Journal of Economics, 116, 969–1007. Skans, O.N. (2005) Age effects in Swedish local labor markets, Economics Letters, 86, 419–26. Staiger, D. and Stock, J.H. (1997) Instrumental Variables Regression with Weak Instruments, Econometrica, 65, 557–86. Stapleton, D.C. and Young, D.J. (1988) Educational Attainment and Cohort Size, Journal of Labor Economics, 6, 330–61. Statistisches Bundesamt (2009) Bevölkerung Deutschlands bis 2060: 12. koordinierte Bevölkerungsvorausberechnung, Statistisches Bundesamt, Wiesbaden. Statistisches Bundesamt (2010) Bevölkerung und Erwerbstätigkeit: Bevölkerung in den Bundesländern, dem früheren Bundesgebiet und den neuen Ländern bis 2060: Ergebnisse der 12. koordinierten Bevölkerungsvorausberechnung, Statistisches Bundesamt, Wiesbaden. vom Berge, P., König, M. and Seth, S. (2013) Sample of Integrated Labour Market Biographies (SIAB) 1975 – 2010, FDZ Datenreport No. 01/2013, Institute for Employment Research, Nuremberg. Welch, F. (1979) Effects of Cohort Size on Earnings: The Baby Boom Babies’ Financial Bust, Journal of Political Economy, 87, S65–S97. Wright, R.E. (1991) Cohort size and earnings in Great Britain, Journal of Population Economics, 4, 295–305. Zimmermann, K. (1991) Ageing and the Labour Market. Age structure, cohort size and unemployment, Journal of Population Economics, 4, 177–200. IAB-Bibliothek 367 94 Regional population structure and young workers’ wages Table S1: Exclusion of observations with more than a full-time job Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.44 (0.68)* -3.29 (1.01)*** -0.81 (0.77) -2.90 (1.22)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 135.19*** 0.32 – – – 0.46 (0.00)*** 131.05*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 102,387 1,296 108 102,387 1,296 108 102,387 1,296 108 102,387 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -1.81%* -4.14%*** -1.03% -3.68%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Figure S3: Histogram of 2SLS baseline coefficients after excluding individual regions Source: Sample of Integrated Employment Biographies (authors’ calculations). The youth-share coefficients are derived from the baseline model; the solid line represents the youth-share coefficients from the full model, the dashed lines the corresponding 95% confidence interval. 0 5 10 15 20 25 30 Frequency -6 -5 -4 -3 -2 -1 0 Youth-share coefficient Place of residence 0 5 10 15 20 25 30 Frequency -6 -5 -4 -3 -2 -1 0 Youth-share coefficient Place of employment 95Chapter 2 Supplementary Material As discussed in the paper, the analysis is restricted to those individuals who are subject to social security contributions. Tables S1 and S2 show the results from further homogenising the sample by either excluding those individuals with more than a full-time job (Table S1) or by dropping individuals whose employment spells contain less than 90 days (Table S2). In the first case the marginal effects are very close to those of the paper’s baseline specification, while they become slightly smaller in the second case. Table S2: Exclusion of observations with employment spells of less than 90 days Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.21 (0.62)†-2.75 (0.97)*** -0.69 (0.71) -2.53 (1.21)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 137.34*** 0.32 – – – 0.46 (0.00)*** 132.74*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 103,652 1,296 108 103,652 1,296 108 103,652 1,296 108 103,652 1,296 108 R20.23 0.23 0.23 0.23 ME(stdev) -1.52%* -3.47%*** -0.88% -3.22%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Since the effect of the size of the youth share may vary between differently educated individuals, the sample is homogenised by first excluding those with tertiary education (Table S3) and then those with tertiary or upper secondary education (Table S4). In neither case is the size of the marginal effects substantially changed. IAB-Bibliothek 367 96 Regional population structure and young workers’ wages Table S3: Exclusion of observations with tertiary education Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.47 (0.62)* -3.22 (0.96)*** -0.91 (0.73) -2.97 (1.22)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 137.01*** 0.32 – – – 0.46 (0.00)*** 132.41*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 106,422 1,296 108 106,422 1,296 108 106,422 1,296 108 106,422 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -1.86%* -4.05%*** -1.16% -3.77%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Table S4: Exclusion of observations with tertiary or upper secondary education Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.39 (0.64)* -3.29 (0.95)*** -0.84 (0.72) -3.04 (1.21)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 - - - 0.46 (0.00)*** 140.36*** 0.33 - - - 0.46 (0.00)*** 135.72*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 101,820 1,296 108 101,820 1,296 108 101,820 1,296 108 101,820 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -1.75%* -4.14%*** -1.06% -3.85%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. 97Chapter 2 Supplementary Material The SIAB dataset contains observations with unrealistically low average daily wages. In order to get a handle on this issue, observations with daily wages below twice the value of the minor-employment threshold were excluded from the sample. Table S5 shows that when these observations are included, the size of the coefficients decrease in size and they become less significant. Table S5: No truncation of the wage distribution Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -0.99 (0.78) -2.51 (1.09)* -0.33 (0.85) -2.14 (1.26)† Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 136.17*** 0.32 – – – 0.46 (0.00)*** 131.61*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 110,651 1,296 108 110,651 1,296 108 110,651 1,296 108 110,651 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -1.25% -3.16%* -0.41% -2.72%† Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. The paper’s empirical analysis is restricted to those individuals who are employed on 30 June of a given year and while June values are usually representative of average annual employment levels, the selection of a specific date is essentially arbitrary. Table S6 shows the results when the reference date is set to 31 December. Doing so produces comparable results in terms of sign and significance but the size of the marginal effects increases in magnitude. IAB-Bibliothek 367 98 Regional population structure and young workers’ wages Table S6: Alternative reference date (31 December) Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.36 (0.65)* -3.80 (0.96)*** -0.75 (0.74) -3.26 (1.16)*** Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 135.93*** 0.32 – – – 0.46 (0.00)*** 130.83*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 113,748 1,296 108 113,748 1,296 108 113,748 1,296 108 113,748 1,296 108 R20.30 0.30 0.30 0.30 ME(stdev) -1.71%* -4.79%*** -0.95% -4.14%*** Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Table S7: Employment-based youth share variable Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -0.45 (0.37) -3.06 (1.08)*** -0.31 (0.41) -2.79 (1.28)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.48 (0.01)*** 56.91*** 0.18 – – – 0.47 (0.01)*** 56.01*** 0.18 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R20.24 0.24 0.24 0.24 ME(stdev) -0.69% -4.72%*** -0.48% -4.30%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. 99Chapter 2 Supplementary Material The youth-share variable is meant to measure the potential supply of young individuals to the labour market. In the paper this variable is constructed from the size of the population in the corresponding age group. However, it is likely that parts of this group are not available to the labour market and as such a population-based variable might provide an inaccurate measure of age-specific labour supply. When a youth-share variable is used instead that is defined as the number of employees aged between 15 and 24 relative to the number of employees between 15 and 64, similarly sized coefficients are obtained, but since the standard deviation of the employment-based youth-share variable is larger than in the case of the population-based variable the marginal effects increase in size (Table S7). The paper estimates the effect of the youth-share on individual wages. Alternatively, it is possible to first average individual-level variables within a region-year cell and to then regress the average daily wage within such a cell on the youth share and weighting the regression by the number of observations in a region-year cell (see Angrist and Pischke, 2009). As can be seen from Table S8, the aggregate-level analysis yields comparable results, though the marginal effects are slightly larger. Table S8: Aggregated model (variables averaged at the level of the region-year cell) Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.65 (0.68)* -3.71 (1.14)*** -0.89 (0.77) -3.37 (1.43)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.45 (0.04)*** 134.84*** 0.30 – – – 0.44 (0.04)*** 124.29*** 0.30 Observations Labour-market regions-year cells Labour-market regions (clusters) 1,296 108 1,296 108 1,296 108 1,296 108 R20.82 0.81 0.82 0.81 ME(stdev) -2.08%* -4.68%*** -1.14% -4.28%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. IAB-Bibliothek 367 100 Regional population structure and young workers’ wages In the empirical specification the effects of age and experience on an individual’s wage are approximated through the inclusion of these variables’ first two polynomials. However, very similar results are obtained if mutually exclusive dummy variables are used instead (Table S9). Table S9: Age and experience dummies Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.47 (0.63)* -3.27 (0.98)*** -0.86 (0.73) -2.96 (1.23)* Dummies Year Labour-market regions Age Experience Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.00)*** 136.64*** 0.32 – – – 0.46 (0.00)*** 131.81*** 0.32 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R20.29 0.29 0.29 0.29 ME(stdev) -1.85%* -4.12%*** -1.10% -3.75%* Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Table S10 contains the results of estimating a double-log specification. The coefficients of the youth-share variable continue to be negative and significant. Comparing the results of Tables 1 and 2 in the paper shows that when a district-specific youth-share variable is used rather than one based on labourmarket regions, the decrease in the size of the coefficient is considerably stronger for the place of employment than the place of residence. In the following, all districts are ordered according to the difference between the number of observations in the sample that live and that work in a district. The model of Equation 1 is then estimated separately for the districts from the top half (i.e. for which the difference is largest) and for those from the bottom half (i.e. for which the difference is smaller) of this ordering. The results are shown in Tables S11 and S12, respectively. As was already discussed in the paper, negative and significant effects are only found for the set of districts from the top half of the 101Chapter 2 Supplementary Material ordering. The row Fraction of full sample shows that for the place-of-residence specification the majority of observations (55%) are from such districts. In the place-of-employment specification the corresponding Figure stands at only 44%. Table S10: Double-log specification Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -0.24 (0.11)* -0.52 (0.15)*** -0.13 (0.13) -0.46 (0.19)* Dummies Year Labour-market regions Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.42 (0.00)*** 129.62*** 0.35 – – – 0.42 (0.00)*** 124.52*** 0.35 Observations Individuals Labour-market regions-year cells Labour-market regions (clusters) 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 107,351 1,296 108 R20.24 0.24 0.24 0.24 Cluster-robust standard errors in parentheses (clustered at the level of the labour-market region). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. Assuming that the districts from the top half represent those which individuals are more likely to live in than to work in, the larger decrease in the size of the youthshare coefficient (i.e. a larger degree of attenuation) in the place-of-employment specification may be explained by the fact that the degree of measurement error is more pronounced in regions that people are more likely to live in and that this type of district is over-represented in the place-of-employment specification. To support this argument, the bottom rows of Tables S11 and S12 show the mean difference between the youth-share variable at the level of the labour-market region and of the district in the corresponding sample. A comparison of these differences between Table S11 and Table S12 shows that regardless of whether the place-of-residence (0.51 as opposed to 0.42) or the place-of-employment specification (0.55 as opposed to 0.40) is used, the extent of measurement error is larger for those districts that individuals are more likely to work in than to live in. IAB-Bibliothek 367 102 Regional population structure and young workers’ wages Table S11: Districts in which individuals are more likely to live Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.78 (0.62)*** -3.63 (0.98)*** -1.07 (0.67) -2.21 (1.12)* Dummies Year Districts Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.46 (0.03)*** 253.16*** 0.45 – – – 0.45 (0.03)*** 230.56*** 0.43 Observations Individuals Fraction of full sample District-year cells Districts (clusters) 58,705 54.69% 1,884 157 58,705 54.69% 1,884 157 47,146 43.92% 1,884 157 47,146 43.92% 1,884 157 R20.23 0.23 0.23 0.23 ME(stdev) -2.12%*** -4.33%*** -1.27% -2.63%* Mean difference 0.42 0.40 Cluster-robust standard errors in parentheses (clustered at the district level). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Mean difference gives the average absolute difference between the district-level youth share and the value at the level of the corresponding labour-market region (multiplied by 100). Finally, Tables S13 and S14 provide the analogues to Tables 3 and 4 but use a youth-share variable that is constructed from districts rather than labour-market regions. First, the results continue to be considerably larger in magnitude for the place of residence than the place of employment when industry and occupation dummies are added; for the place of employment, none of the coefficients are statistically significant. Second, the youth-share coefficients of the districtspecific model remain smaller than the ones from the region-specific model. In the case of the place of residence the district-specific coefficients are smaller by between 27% (industry dummies) and 10% (occupation dummies). In contrast, the differences in size are much more pronounced at the place of employment where the inclusion of dummies for an individual’s industrial or occupational affiliation further reduces the magnitude of the youth-share coefficients relative to those from the labour-market specification. This finding illustrates that the distinction between place of employment and place of residence is of particular importance for the estimated size and significance of the effects at the district level. 103Chapter 2 Supplementary Material Table S12: Districts in which individuals are more likely to work Dependent variable: log real daily earnings Place of residence Place of employment OLS 2SLS OLS 2SLS Youth share -1.05 (0.61)†-1.75 (1.39) 0.11 (0.52) -0.87 (1.43) Dummies Year Districts Control variables Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes First-stage statistics Instrument First-stage statistics F-statistic Shea’s partial R2 – – – 0.43 (0.05)*** 85.01*** 0.17 – – – 0.42 (0.05)*** 60.10*** 0.15 Observations Individuals Fraction of full sample District-year cells Districts (clusters) 48,646 45.31% 1,872 156 48,646 45.31% 1,872 156 60,205 56.08% 1,872 156 60,205 56.08% 1,872 156 R20.27 0.27 0.29 0.29 ME(stdev) -1.63%†-2.70% 0.18% -1.40% Mean difference 0.51 0.55 Cluster-robust standard errors in parentheses (clustered at the district level). ***/**/*/† indicate significance at the 0.005/0.01/0.05/0.10 level, respectively. Instrument shows the coefficient of the instrument in the first-stage regression. ME(stdev) gives the percentage change in daily earnings given an increase in the youth share by one standard deviation. Mean difference gives the average absolute difference between the district-level youth share and the value at the level of the corresponding labour-market region (multiplied by 100). IAB-Bibliothek 367 110 Cohort size and youth labour-market outcomes: the role of measurement error Shimer (2001) provides a theoretical foundation to his empirical findings in the form of a search and matching model with on-the-job search. Changes in the size of the youth population tend to be predictable, as evidenced by the explanatory power of lagged birth rates for the size of the current youth share. Moreover, young individuals are more often either without a job or less well matched than older individuals and are therefore, on average, more willing to take up or switch jobs. This makes it easier for firms to make a productive match with workers in markets with a large number of potential employees. They therefore react to an expected change in the youth share by creating vacancies, to the benefit of all age groups. Aiming to explain the substantial differences between his own and Korenman and Neumark’s (2000) empirical findings, Shimer (2001) points out that the former ignored the possibility of changes in the youth share having an effect on the unemployment rate of other age groups. Specifically, Korenman and Neumark’s (2000) model includes the adult unemployment rate, alongside the youth share, as a regressor in the model of the youth unemployment rate. According to Shimer (2001), if changes in the youth share affect the unemployment rates of both age groups, the former’s coefficient will be biased upwards and he is able to show this using his own dataset. However, applying his empirical model to the data of Korenman and Neumark (2000) produces inconclusive results, which casts doubt on the applicability of his theoretical model to other countries and time periods. The small number of studies that have since looked at the relationship between age structures and unemployment outcomes have yielded mixed results. Using data on Swedish labour markets for the years 1985–1999, Skans (2005) finds no evidence for an effect of the relative size of the group aged 16–24 on the total unemployment rate, but his results are otherwise in line with Shimer (2001) since they show that the youth unemployment rate falls when the size of young age groups increases. In contrast, Foote (2007) shows that when the time dimension of Shimer’s (2001) dataset is extended to 2005 the negative effect of the youth share on the overall unemployment rate decreases considerably and becomes insignificant in most specifications. The empirical evidence of Biagi and Lucifora (2008) also contradicts the findings of Shimer (2001): their analysis of a dataset of European countries spanning the late 1970s to the early 2000s suggests that the share of individuals aged 15–24 has a positive effect on the unemployment rate of the young and is not statistically significant for the unemployment outcomes of prime-age individuals. Finally, Garloff et al. (2013), using data on West German labour-market regions for the years 1993–2008, find that increases in the share of individuals aged 15–24 years are associated with increases in the overall unemployment rate. 111 Empirical analysis Chapter 3 In light of the conflicting results produced by previous studies this analysis provides new evidence on the relationship between age-group size and agespecific unemployment outcomes. Our dataset is a longitudinal sample of European regions covering 2005–2012 which provides us with more heterogeneity to separate the effects of cohort size from other influences than has generally been available in the literature. However, the paper’s main contribution is to consider the effect of the definition of the youth population on the estimates obtained. The previous literature has used the share of individuals aged either 15–24 or 16–24 as a definition of the youth share. Since a high proportion of this group will be in education and therefore potentially unavailable to the labour market, this will, as discussed in the introduction and in more detail below, have important implications for both the interpretation and econometric identification of the cohort-size effect. 3 Empirical analysis 3.1 Data The major part of the dataset that is used in the empirical analysis is constructed by combining different longitudinal EU-SILC releases.1 Appending data from different releases not only allows the extension of the sample period beyond the four years provided by a single longitudinal release, but also increases the number of observations within a given year. In order to match observations from different releases that refer to the same individual, a unique personal identifier is constructed.2 This is then used to verify that there are very few individuals with inconsistencies in age and sex over time3 (see Moffat and Roth, 2016, for further details on the process of appending the different datasets and Berger and Schaffner, 2015, for general information about EU-SILC). Individuals in EU-SILC are not randomly sampled and weights are therefore provided so that unbiased population estimates may be calculated. We use these to construct two new weighting variables: the first of these variables corrects 1 The longitudinal releases are: 2013 (version 1 from 01-08-2015), 2012 (version 3 from 01-08-2015), 2011 (version 4 from 01-03-2015), 2010 (version 5 from 01-08-2014), 2009 (version 4 from 01-03-2013), 2008 (version 4 from 01-03-2012), 2007 (version 5 from 01-08-2011), 2006 (version 2 from 01-03-2009) and 2005 (version 1 from 15-09-07). 2 This identifier is defined as a combination of an observation’s identification number (which is not unique across countries), his country of residence and the rotational group to which he belongs. 3 In total, there are 36 individuals (182 observations) with inconsistencies. All of these individuals are from France, Luxembourg or Norway (i.e. countries in which individuals can be followed for more than 4 years). For these individuals, the inconsistent observations are dropped. If there are only two observations per individual, both are dropped. IAB-Bibliothek 367 112 Cohort size and youth labour-market outcomes: the role of measurement error the initial weights for the number of rotational groups within a country-year combination that change as a result of appending data from different releases (see Moffat and Roth, 2016). The second weighting variable also re-scales the weights so that the size of the estimated population within a region-year-age-sex cell is identical to the statistics reported by Eurostat.4 The so-constructed dataset contains 2.76 million observations on just over 1 million individuals and covers the years 2004–2013. In addition to the country that an individual resides in, EU-SILC provides information about the region of residence at the first level of the Nomenclature of Territorial Units in Statistics (NUTS). Availability of this information allows us to construct the relevant variables at the regional rather than at the national level, which is attractive because estimates of functional labour markets have tended to show them to be defined at the sub-national level (see Moffat and Roth, 2016). Rather than focussing on outcomes at the individual level, the empirical analysis in this paper is concerned with estimating the effect of age-specific cohort size on unemployment and employment outcomes at the level of the corresponding age group. For this reason, the dataset is aggregated to the level of region-yearage cells. The resulting dataset is further supplemented by variables taken from Eurostat’s publicly available database5: the level of regional GDP and the size of relevant age groups between 1991 and 1998 which are used as instruments in the empirical analysis.6 Due to data limitations, observations from the following countries are dropped: Germany, the Netherlands and Portugal (information on NUTS1 regions is not provided); Croatia (lagged population data for the construction of the instrument is not available); Finland, Iceland and Slovenia (age-related variables are randomly perturbed to prevent disclosure); Ireland and the United Kingdom (the age variable is measured at a different time of year for these countries, see footnote 4). Moreover, we exclude observations from Bulgaria, Cyprus, Malta, Norway and Romania because the necessary variables are not available throughout the whole sample period. This leaves a panel of 49 NUTS1 regions from the following countries for which age groups can be observed from 2005–2012 (number of regions per country in parentheses): Austria (3), Belgium (3), Czech Republic (1), 4 Note that while the Eurostat statistics refer to 1 January of a given year, use of the variable age at the end of the income reference period ensures that the population sizes estimated from EU-SILC data refer to 31 December of the preceding year. 5 The data can be obtained through the following link: http://epp.eurostat.ec.europa.eu/portal/page/portal/statistics/ search_database 6 Due to a change in delineation lagged population data is not available before the year 2003 for the two regions ITH (Northeast Italy) and ITI (Central Italy). Since these changes are minor compared to the total size of the regions we instead use lagged age-group size based on the predecessor regions ITD and ITE, which we obtain from the homepage of the Italian Statistical Office (www.istat.it). 113 Empirical analysis Chapter 3 Denmark (1), Estonia (1), Greece (4), Spain (7), France (8), Hungary (3), Italy (5), Lithuania (1), Luxemburg (1), Latvia (1), Poland (6), Sweden (3), Slovakia (1). 3.2 Variables and sample This section serves several purposes: first, it defines the main variables of the empirical model; second, it discusses the age range of the sample; finally, an illustration is provided of the variation in the cohort-size variable that is used for identification. The analysis separately estimates the effect of changes in cohort size on the share of individuals in age group j, region r and year t that are unemployed (unempjrt ) and employed (empjrt ). As discussed in the previous section, these shares are derived from individual-level data. Specifically, the weighted sum of male individuals who report to be (un-)employed in a given region-year-age group is calculated and divided by the total male population in that cell. Female observations are excluded in order to avoid the results being affected by selected labour-market participation. As these variables are standardised on the population rather than the labour force, the outcome variables differ from the unemployment and the employment rate. An advantage of this specification is that any effects that changes in cohort size, if measured without error, might have on participation rates could be ignored in the interpretation of the results. Figure A1 in the Appendix shows the development of the dependent variables unempjrt and empjrt as well as of a similarly defined variable that shows the share of individuals reporting to be in education in a given age group (educjrt ). These variables are plotted for the age group 18–29 in selected regions and years to illustrate the variation in age-specific labour-market outcomes across Europe. While there are differences in the slope of the profiles, a common feature of all region-year combinations is that the employment share tends to increase and the share of individuals in education decreases with age. In contrast, there is no obvious trend in the unemployment share. In order to understand the implications of the high share of young individuals in education, the empirical model is firstly estimated for overlapping five-year age groups (beginning with individuals aged 18–22 and ending with individuals aged 25–29). The reason for adopting this strategy is that for younger age groups the coefficients will capture the effect of cohort size on labour market participation and, conditional on participation, the effect on (un-)employment. If the decision to participate in the labour market is also affected by cohort size, the estimated effects on employment and unemployment would be confounded by the effect of cohort size on participation. Moreover, the existence of measurement error in the cohort-size variable among IAB-Bibliothek 367 114 Cohort size and youth labour-market outcomes: the role of measurement error young age groups, as described further in Section 3.3, may also lead to biased estimates. We therefore focus on individuals aged 25–29 since the estimates for this group will be less susceptible to these problems since, as shown in Figure A1, the share of individuals in education has decreased substantially by that age. Means and standard deviations of the three dependent variables are shown in the first two columns of Table 1 for the age range 25–29. On average 78% of individuals in a region-year-age group cell are employed compared to 13% that are unemployed. The three remaining columns provide an insight into whether these variables tend to vary most across regions, years or age groups. This is done by regressing each of the dependent variables on a set of dummy variables for two of the aforementioned dimensions and then comparing the adjusted R2. Dummies for years and age groups explain only 14% of the variation in the employment share but this value increases considerably once region dummies are included, which suggests that most of the variation in this variable exists between regions. While the explanatory power of the dummy variables is generally lower, the between-region variation also appears to be largest for the unemployment share. Table 1: Descriptive statistics (employment and unemployment share) Mean Standard deviation Adjusted R2 (year, age) Adjusted R2 (region, age) Adjusted R2 (region, year) Empjrt 0.777 0.156 0.136 0.459 0.394 Unempjrt 0.126 0.109 0.063 0.281 0.333 Means and standard deviations are weighted by the weight-adjusted number of individuals per region-yearage group cell. Adjusted R2 is derived from a regression of the dependent variables on dummies for the indicated variables; the regression is weighted by the weight-adjusted number of individuals per region-year-age group cell. The main explanatory variable measures age-specific cohort size which refers to the number of individuals in age group j, region r and year t, Njrt , relative to the size of the population aged between 16 and 65, N16–65, rt . While most studies instead use a measure of the youth share, e.g. the relative size of the age group 16–24, we choose a specification that also varies across age to better capture the assumption of imperfect substitutability across age groups which has been posited in theoretical models (Card and Lemieux, 2001). Since it seems overly restrictive to assume that individuals only compete with individuals of the same age, we adopt another specification that has been previously used in this literature (Wright, 1991; Brunello, 2010).7 This defines the cohort-size variable as a weighted sum 7 We show in the Supplementary Material that alternative specifications of the cohort-size variable, including unweighted sums across three and five age groups, yield comparable results to those shown in Table 3. 115 Empirical analysis Chapter 3 that takes into account the size of the age groups that are up to two years older or younger than the reference group as shown in Equation 1: CSjrt = (1/9)Nj – 2, rt + (2/9)Nj – 1, rt + (3/9)Njrt + (2/9)Nj + 1 , rt + (1/9)Nj + 2, rt N16 – 65, rt [1] These quantities are estimated from the EU-SILC dataset by computing the weighted sum of male and female observations in the corresponding region-yearage cells. As they are not available to the labour market, individuals reporting to be either in the military or disabled or unfit to work are omitted but individuals reporting that they are in education are included (the implications of this are discussed in Section 3.3). The size of an age group in a given region and year is not necessarily exogenous because individuals might react to contemporaneous economic shocks by migrating into regions that offer better economic prospects. If such selfselection takes place, cohort-size would be endogenous to the share of individuals that are (un-)employed and estimation by ordinary least squares (OLS) would yield an inconsistent estimate of the cohort-size effect. We address this issue by employing an IV strategy in which the cohort size of the age group that is fourteen years younger than the reference group as observed fourteen years earlier serves as an instrument. Identification strategies based on time-lagged and age-lagged instruments or, as a special case of the former, birth rates are common in this literature (Korenman and Neumark, 2000; Shimer, 2001; Skans, 2005; Garloff et al., 2013; Moffat and Roth, 2016).8 Instruments of this type are appealing because a cohort that was relatively large (small) in the past is likely to remain large (small) in the present despite migration and natural population changes9: CS_Insjrt = (1/9)Nj – 16, r, t – 14 + (2/9)Nj – 15, r, t – 14 + (3/9)Nj – 14r, t – 14 + (2/9)Nj – 13, r, t – 14 + (1/9)Nj – 12, r, t – 14 N 2 – 51, r, t – 14 [2] 8 If cohort-size effects are heterogeneous across age, region and/or time, 2SLS estimates a local average treatment effect (LATE) (Imbens and Angrist, 1994). This estimate is the weighted average of the region-year-age cell-specific effects of cohort size with the largest weights attached to cells for which the relationship between the instrument and cohort-size is strongest (Angrist and Imbens, 1995). Since the strength of the relationship between the instrument and cohort-size will be mainly determined by net migration, greater weight will be attached to cells with low levels of net migration. If immigrants are less attractive to employers as a result of having less countryspecific human capital (Kim and Park, 2013) than individuals that lived in the region fourteen years ago, this suggests that the LATE will be more positive (more negative) in the employment (unemployment) model than the average treatment effect (ATE). 2SLS estimates may then be larger than OLS estimates of the cohort-size effects if this effect outweighs that of self-selection bias, which would tend to cause OLS to overestimate the positive (negative) effect on employment (unemployment). 9 Further information on the instrument can be found in Moffat and Roth (2016), while the validity of timeand age-lagged instruments is discussed in Garloff and Roth (2016). IAB-Bibliothek 367 116 Cohort size and youth labour-market outcomes: the role of measurement error Table 2 contains descriptive statistics on the cohort-size variable and its instrument. On average, the five-year weighted sum of an age group in the range 25–29 accounts for about 2% of the population aged between 16 and 65, while the value is slightly smaller in the case of the instrument. For both variables, the larger part of the variation exists between regions. Table 2: Descriptive statistics (cohort-size variable and instrument) Mean Standard deviation Adjusted R2 (year, age) Adjusted R2 (region, age) Adjusted R2 (region, year) CSjrt 0.021 0.003 0.073 0.749 0.778 CS_Insjrt 0.020 0.003 0.080 0.780 0.826 Means and standard deviations are weighted by the weight-adjusted number of individuals per region-yearage group cell. Adjusted R2 is derived from a regression of the dependent variables on dummies for the indicated variables; the regression is weighted by the weight-adjusted number of individuals per region-year-age group cell. Figures A2 and A3 plot the dependent variables and the cohort-size variable (depicted as the fitted value from a weighted regression on the instrument) across time and age groups, respectively, for the same set of regions as in Figure A1 and thereby illustrate the variation from which cohort-size effects can be identified. Variation over time for given combinations of regions and age groups can be seen in Figure A2; the chosen regions are representative of the larger parts of Europe to which they belong: in Western and Northern Europe (represented by regions BE2 and SE1), the cohort-size profiles are rather flat. In contrast, in region ES5 there is a clear decrease in cohort size over time which affects all age groups – similar profiles can be found in the remaining regions of Spain as well as in Greece and Italy. Finally, different types of profiles can be found in Eastern Europe: on the one hand, the decreasing trend in cohort size in region HU1 resembles the developments in Southern Europe, while on the other hand age groups have increased in size in the Baltic country Latvia. Figure A3 suggests that variation across age groups is less pronounced: older age groups tend to be larger in ES5 and HU1, but the differences become smaller in later years. The profiles in the remaining regions are comparatively flat. At the same time both figures also illustrate the variation in cohort size across regions for given years and age groups. For example, the share of older age groups is larger in regions ES5 and HU1 in earlier years, whereas younger cohorts are relatively big in LV0 at the end of the sample period. While the regression analysis in Section 4 makes use of variation across each of these dimensions, in the Appendix we show results that are obtained from a single source of variation. 117 Empirical analysis Chapter 3 3.3 Model According to the theory outlined in the literature review, age-specific labour market outcomes are determined by the supply of age-specific labour. Therefore the effect of cohort size on the outcome variables is modelled as shown in Equation 3 where sharejrt represents either the unemployment or employment share, CS* jrt represents measurement error-free cohort size (i.e. the size of the age cohort that is available to the labour market), xjrt represents a vector of control variables and ε jrt is an error term: sharejrt = α + β CS* jrt + x' jrt γ + ε jrt [3] In addition to the problem of regional self-selection that is addressed by IV estimation, there is also a problem of measurement error. This has so far not been addressed in this literature. It arises because of the inclusion of individuals, many of whom will be in education, that are unavailable to the labour market in the cohort-size variable. Moreover, datasets usually do not allow distinguishing individuals that are committed to long-term educational programmes and therefore unavailable to the labour market from individuals in education that would enter the labour market if an attractive opportunity arose (Jones and Riddell, 2006; Moffat and Yoo, 2015). The existence of the latter group means that the alternative approach of excluding those in education from the cohortsize variable would not provide a solution to the measurement-error problem.10 Formally, the relationship between the observable age-specific cohort-size variable CSjrt and the unobservable measurement error-free variable can be represented as follows: CSjrt = CS* jrt + ujrt [4] In Equation (4), ujrt is the part of observed cohort size that is not available to the labour market (i.e. the measurement error). Rearranging and substituting Equation (4) into Equation (3) gives: sharejrt = α + β CSjrt + x' jrt γ + ε jrt – β ujrt [5] 10 In the Supplementary Material we provide the regression results from a model in which the numerator of the cohort-size variable is constructed from individuals reporting to be employed or unemployed. For the age group 25–29 the obtained results are very similar to those reported in Table 3. Using younger age groups produces a pattern of cohort-size coefficients which is close to the one in Figure 1 which suggests that exclusion of those reporting to be in education does not remove the problem of measurement error. IAB-Bibliothek 367 118 Cohort size and youth labour-market outcomes: the role of measurement error If the measurement error is ‘classical’, there is no correlation between the error-free measure of cohort size and the measurement error and this leads to attenuation of the estimated effect of cohort size. However, empirical evidence suggests that members of large cohorts are less likely to acquire education (Fertig et al., 2009), which suggests the existence of a correlation between the size of an age group CSjrt and ujrt. Arguably, the number of individuals who are available to the labour market is larger in larger age groups and therefore the correlation between the degree of measurement error and the observable cohort size also carries over to the latent variable CS* jrt , which measures the size of an age group that is available to the labour market. In this ‘non-classical’ case, it is not possible to state a priori the direction of bias since it will be dependent on the relative variances of CS* jrt and ujrt , the size of the covariance of CS* jrt and ujrt and the partial correlations between the measurement error and the dummy variables in the model (Bound et al., 2001). A second reason for the existence of non-classical measurement error is given by the current demographic processes, as a result of which younger age groups tend to be smaller than older ones in a given region and year (support for this hypothesis is provided in the Supplementary Material). Moreover, given the assumption that the share of non-participants is larger in younger age groups – for which the substantially larger education shares in younger age groups provide some evidence – it is possible for the latent cohort-size variable and the degree of measurement error to be negatively correlated across age groups. This will be the case as long as the ratio of the non-participation share in younger and older groups exceeds the ratio of the size of older and younger groups (details on this argument are provided in the Supplementary Material). While two-stage least squares (2SLS) estimation is one approach to tackling measurement error (Hausman, 2001), the instrument which is standard in the literature does not purge the correlation with ujrt . The instrument is based on the size of the same cohort observed at an earlier point in time and since an age group that is relatively large in the present can be expected to have also been relatively large in the past, the instrument would also be correlated with the degree of measurement error. As a result, 2SLS will not provide a consistent estimate of the cohort-size effect. For the sample of individuals aged 25–29, the empirical analysis is based on 1,959 region-year-age cells11. Two specifications of Equation 5 are estimated for each of the outcome variables. Analogously to the use of control variables in Shimer (2001), in the baseline specification vector xjrt only contains a constant 11 In principle, 5 age groups (25–29) are observed in 49 regions for 8 years (2005–2012), but since there are no observations for age group 26 in region FR1 and year 2010 in the sample, the total number of observations is reduced by one. 119 Results Chapter 3 and three sets of dummy variables for each of the three dimensions of the cohortsize variable: regions, years and age groups. In the second specification a set of control variables is added to the model (definitions and summary statistics are given in Table A1 in the Appendix). One part of these variables is assumed to affect the (un-)employment probability at the individual level and has therefore been aggregated in order to control for compositional differences between regionyear-age cells. They include the share of individuals in such cells that a) belong to different educational groups according to the International Standard Classification of Education (ISCED), b) are married and c) reside in areas that differ with respect to their degree of urbanisation. Moreover, we add the level of regional GDP. While the use of year dummies accounts for shocks that are common to all region-age cells, this variable is useful in order to control for the region-specific economic environment in a given year. The inclusion of regional GDP therefore helps to avoid the estimated cohort-size effects being confounded by regional economic shocks. 4 Results Figure 1 shows the estimated coefficients and confidence intervals on the cohortsize variable using overlapping samples of differently aged individuals when the dependent variable is the unemployment and employment share, respectively. For both outcome variables, the effect of cohort size varies substantially across age groups. When the dependent variable is the unemployment share, the effects are positive and statistically significant for individuals aged 18–22 but are negative and statistically significant for older groups. The effect appears to converge to between -10 and -20 for the older groups. The shift in sign and magnitude of the coefficients coincides with a decrease in the share of individuals reporting to be in education (see Figure A1 in the Appendix). In the employment model, cohortsize effects are significant and negative for individuals aged 18–22 but positive and significant for older age groups, converging to a value of approximately 25. The results for the younger age groups appear to be supportive of the cohortcrowding hypothesis. However, our view is that the estimated effects for younger age groups cannot be regarded as a direct test of this hypothesis since they capture both the effect of cohort size on labour-market participation and the effect on (un-)employment. For example, the finding that cohort size reduces the employment share of individuals aged 18–22 may indicate either that large cohorts lead young individuals to acquire education and thereby defer entry to the labour market or that young individuals in the labour market are disadvantaged by belonging to a large age group. In addition to this problem of interpretation, the change in the coefficients may be driven by measurement error in the cohort-size IAB-Bibliothek 367 126 Cohort size and youth labour-market outcomes: the role of measurement error Appendix Figure A1: Development of employment, unemployment and education shares across age groups Source: EU-SILC (authors’ calculations). BE2: the Flemish region of Belgium; ES5: East Spain; HU1: Central Hungary; LV0: Latvia, SE1: East Sweden. 0101010101 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 BE2, 2005 BE2, 2007 BE2, 2009 BE2, 2011 ES5, 2005 ES5, 2007 ES5, 2009 ES5, 2011 HU1, 2005 HU1, 2007 HU1, 2009 HU1, 2011 LV0, 2005 LV0, 2007 LV0, 2009 LV0, 2011 SE1, 2005 SE1, 2007 SE1, 2009 SE1, 2011 Emp Unemp Educ Education and (un-)employment shares Age 127 Appendix Chapter 3 Figure A2: Development of employment and unemployment shares and of fitted cohort-size variable over time Source: EU-SILC (authors’ calculations). BE2: the Flemish region of Belgium; ES5: East Spain; HU1: Central Hungary; LV0: Latvia, SE1: East Sweden. 0.015 0.0250.015 0.0250.015 0.0250.015 0.0250.015 0.025 0101010101 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 25 26 27 28 29 ES5, 2005 ES5, 2007 ES5, 2009 ES5, 2011 HU1, 2005 HU1, 2007 HU1, 2009 HU1, 2011 LV0, 2005 LV0, 2007 LV0, 2009 LV0, 2011 SE1, 2005 SE1, 2007 SE1, 2009 SE1, 2011 BE2, 2005 BE2, 2007 BE2, 2009 BE2, 2011 (Un-)employment share Cohort size Age Emp Unemp Cohort size (fitted values) IAB-Bibliothek 367 128 Cohort size and youth labour-market outcomes: the role of measurement error Figure A3: Development of employment and unemployment shares and of fitted cohort-size variable over age groups Source: EU-SILC (authors’ calculations). BE2: the Flemish region of Belgium; ES5: East Spain; HU1: Central Hungary; LV0: Latvia, SE1: East Sweden. 0101010101 ES5, 25 ES5, 27 ES5, 28 ES5, 29 HU1, 25 HU1, 27 HU1, 28 HU1, 29 LV0, 25 LV0, 27 LV0, 28 LV0, 29 SE1, 25 SE1, 27 SE1, 28 SE1, 29 BE2, 25 BE2, 27 BE2, 28 BE2, 29 (Un-)employment share 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 2005 2007 2009 2011 0.015 0.0250.015 0.0250.015 0.0250.015 0.0250.015 0.025 Cohort size Age Emp Unemp Cohort size (fitted values) 129 Appendix Chapter 3 Table A1: Definitions and descriptive statistics of control variables Name Definition Source Mean Standard deviation ISCED_0 Share of individuals in region-year-age cell with pre-primary education EU-SILC 0.006 0.027 ISCED_1 Share of individuals in region-year-age cell with primary education EU-SILC 0.040 0.060 ISCED_2 Share of individuals in region-year-age cell with lower secondary education EU-SILC 0.136 0.133 ISCED_3 Share of individuals in region-year-age cell with upper secondary education EU-SILC 0.479 0.187 ISCED_4 Share of individuals in region-year-age cell with post-secondary, non-tertiary education EU-SILC 0.035 0.052 ISCED_5 Share of individuals in region-year-age cell with tertiary education (also includes category ISCED_6, i.e. individuals with second stage of tertiary education) EU-SILC 0.304 0.168 Married Share of individuals in region-year-age cell that are married EU-SILC 0.195 0.153 Urban_1 Share of individuals in region-year-age cell living in densely populated areas (an area with a population density of more than 500 inhabitants per square kilometre (km) and a population of at least 50,000 inhabitants) EU-SILC 0.461 0.216 Urban_2 Share of individuals in region-year-age cell living in intermediately populated areas (an area with a population density of more than 100 inhabitants per square km and either a population of at least 50,000 inhabitants or adjacent to a ‘densely populated’ area) EU-SILC 0.248 0.170 Urban_3 Share of individuals in region-year-age cell living in thinly populated areas (an area with fewer than 100 inhabitants per square km and a population of less than 50,000 inhabitants) EU-SILC 0.291 0.222 GDP Gross domestic product at the NUTS1 level (in billion Euros, adjusted for purchasing-power-parity) Eurostat 188.391 127.737 Means and standard deviations are weighted by the weight-adjusted number of individuals per region-yearage group cell. IAB-Bibliothek 367 130 Cohort size and youth labour-market outcomes: the role of measurement error Table A2: Full OLS and 2SLS regression results (Unemployment share) Unemployment share OLS 2SLS OLS 2SLS Cohort size -10.32*** (1.70) -17.30*** (2.10) -7.98*** (1.73) -15.06*** (2.05) Dummies Region Year Age Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Control variables ISCED_1 ISCED_2 ISCED_3 ISCED_4 ISCED_5 Married Urban_2 Urban_3 GDP – – – – – – – – – – – – – – – – – – 0.07 (0.14) 0.06 (0.13) -0.01 (0.12) -0.11 (0.13) -0.08 (0.13) -0.10*** (0.02) -0.03 (0.03) -0.03 (0.03) -0.00*** (0.00) 0.08 (0.14) 0.06 (0.13) -0.01 (0.12) -0.09 (0.13) -0.08 (0.13) -0.09*** (0.02) -0.03 (0.03) -0.03 (0.03) -0.00*** (0.00) Observations Region-year-age cells 1,959 1,959 1,959 1,959 R20.38 0.37 0.41 0.40 F-stat – 1,540.67*** – 1,642.59*** ME(std) -0.03*** -0.05*** -0.02*** -0.05*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. 131 Appendix Chapter 3 Table A3: Full OLS and 2SLS regression results (Employment share) Employment share OLS 2SLS OLS 2SLS Cohort size 14.39*** (2.03) 24.32*** (2.64) 11.91*** (2.02) 22.07*** (2.52) Dummies Region Year Age Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Control variables ISCED_1 ISCED_2 ISCED_3 ISCED_4 ISCED_5 Married Urban_2 Urban_3 GDP – – – – – – – – – – – – – – – – – – 0.47** (0.20) 0.51** (0.20) 0.57*** (0.19) 0.66*** (0.20) 0.62*** (0.19) 0.09*** (0.03) 0.06* (0.03) 0.07* (0.04) 0.00*** (0.00) 0.46** (0.20) 0.51** (0.20) 0.58*** (0.19) 0.64*** (0.20) 0.62*** (0.19) 0.09** (0.03) 0.06* (0.03) 0.07* (0.04) 0.00*** (0.00) Observations Region-year-age cells 1,959 1,959 1,959 1,959 R20.53 0.52 0.56 0.55 F-stat – 1,540.67*** – 1,642.59*** ME(std) 0.04*** 0.08*** 0.04*** 0.07*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. IAB-Bibliothek 367 132 Cohort size and youth labour-market outcomes: the role of measurement error Table A4: First-stage regression results Unemployment share Employment share Instrument 0.93*** (0.02) 0.93*** (0.02) 0.93*** (0.02) 0.93*** (0.02) Dummies Region Year Age Control variables Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes No Yes Yes Yes Yes Observations Region-year-age cells 1,959 1,959 1,959 1,959 R20.92 0.92 0.92 0.92 F-stat 1,540.67*** 1,642.59*** 1,540.67*** 1,642.59*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. 133 Appendix Chapter 3 Table A5: OLS and 2SLS results Panel A: Unemployment share OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -12.84*** (1.91) -23.01*** (2.37) -10.76*** (1.67) -17.35*** (2.07) -2.16 (2.15) -1.46 (2.64) Dummies Region Year Age Region-by-age Year-by-age Region-by-year Control variables Yes Yes Yes Yes No No No Yes Yes Yes Yes No No No Yes Yes Yes No Yes No No Yes Yes Yes No Yes No No Yes Yes Yes No No Yes No Yes Yes Yes No No Yes No Observations Region-year-age cells 1,959 1,959 1,959 1,959 1,959 1,959 R20.43 0.41 0.39 0.38 0.57 0.57 F-stat – 1,140.11*** – 1,582.57*** – 674.72*** ME(std) -0.04*** -0.07*** -0.03*** -0.05*** -0.01 -0.00 Panel B: Employment share OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 13.88*** (2.24) 26.55*** (2.89) 14.41*** (2.02) 24.40*** (2.59) 7.24*** (2.73) 11.15*** (3.37) Dummies Region Year Age Region-by-age Year-by-age Region-by-year Control variables Yes Yes Yes Yes No No No Yes Yes Yes Yes No No No Yes Yes Yes No Yes No No Yes Yes Yes No Yes No No Yes Yes Yes No No Yes No Yes Yes Yes No No Yes No Observations Region-year-age cells 1,959 1,959 1,959 1,959 1,959 1,959 R20.58 0.57 0.54 0.53 0.66 0.66 F-stat – 1,140.11*** – 1,582.57 – 674.72*** ME(std) 0.04*** 0.08*** 0.04*** 0.08*** 0.02*** 0.03*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. IAB-Bibliothek 367 134 Cohort size and youth labour-market outcomes: the role of measurement error Supplementary material S1 Selection of the age range and measurement error The paper’s main finding is that the estimated effect of cohort size on the (un-) employment share is sensitive to the selected age range of the sample (see Figure 1 in the paper). We propose two explanations for the observed pattern of the coefficients and in both cases the core of the argument is that for young age groups the cohort-size variable can be a poor measure of the age-specific supply of labour: first, a population-based cohort-size variable will include a substantial number of individuals that are not on the labour market, primarily because they are acquiring education; second, given the large share of non-participants among young age groups the estimated effect of cohort-size on the (un-)employment share will be confounded by the former’s effect on the decision to participate in the labour market. In the following, we provide further detail on the former point. Figures S1 and S2 plot the share of individuals reporting to be in education against age for different region-year combinations.12 As can be seen, the education share can be close to 100% at age 18 and usually is in excess of 50% at age 20, whereas the share is considerably smaller in the age range 25–29, which is used in the empirical analysis of this paper.13 This observation provides support for the hypothesis that the share of individuals that are included in a population-based cohort-size variable but that are not on the labour market can be substantial, especially among young age groups. However, it is important to note that simply excluding those individuals that report to be in education from the construction of the cohort-size variable does not necessarily lead to a better measure of agespecific labour supply. First, a part of the group of individuals reporting to be in education may be enticed to enter the labour market depending on the conditions of employment and as such should be treated as being available to the labour market, whereas participants in lengthy degree programmes are less likely to do so (these groups cannot be separated in the data); second, switching between periods of participation and non-participation is more likely to occur among young individuals compared to older age groups whose members tend to be more established in the labour market. 12 The regions are EL3 (Attica), ES3 (Madrid), ES6 (Andalusia), ITF (Southern Italy) and ITH (Northeast Italy), CZ0 (Czech Republic), DKO (Denmark), FR1 (Île de France), LT0 (Lithuania) and PL1 (Central Poland). 13 The main exception is Denmark where the education share takes longer to decrease and can be large at later ages (e.g. age 26 in the year 2011). However, we are able to show in Figures S3 and S4 that the exclusion of Denmark from the sample has virtually no effect on the size of the coefficient in the unemployment and the employment model, respectively, while allowing the sample to start at age 26 instead of 25 also yields comparable coefficients in both models (see Figure S6). 135 Supplementary material Chapter 3 Figure S1: Development of education share and fitted cohort-size variable (set 1) Source: EU-SILC (authors’ calculations). 0101010101 ES3, 2005 ES3, 2007 ES3, 2009 ES3, 2011 ES6, 2005 ES6, 2007 ES6, 2009 v, 2011 ITF, 2005 ITF, 2007 ITF, 2009 ITF, 2011 ITH, 2005 ITH, 2007 ITH, 2009 ITH, 2011 EL3, 2005 EL3, 2007 EL3, 2009 EL3, 2011 Education share 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 18 20 22 24 26 28 .5.5 .5.5 .5 0.015 0.0250.015 0.0250.015 0.0250.015 0.0250.015 0.025 Educ Cohort size (fitted values) Age IAB-Bibliothek 367 142 Cohort size and youth labour-market outcomes: the role of measurement error The effects of excluding individual age groups from the sample are illustrated in Figure S6. The largest change in the coefficient can be observed when the age group 29 is dropped, in which case the magnitude of the coefficient increases in both models and almost moves outside of the full sample’s confidence interval in the employment model. The responsiveness of labour-market shares to changes in cohort size therefore appears less pronounced for this age group. Unfortunately, the unavailability of lagged population data prevents the inclusion of older age groups in the sample and thus the possibility to check whether a further decrease in the strength of the relationship between cohort size and labour-market shares could be found at older ages. Such a development would be in line with the underlying mechanism that is proposed by Shimer (2001): firms create vacancies in areas where the share of young individuals is large because the former are usually not well matched to their jobs and a large pool of such individuals makes it easier for firms to find good matches for these vacancies. However, if the degree to which individuals are matched to their job increases with age, larger older age groups would not necessarily induce the same reaction on the firms’ side because members of those age groups would not be as easily enticed to engage in on-the-job search as younger individuals, thereby reducing the incentive to firms to create vacancies. In addition, dropping age 25 also increases the magnitude of the coefficient in the unemployment model but has no sizeable effect in the employment model. Figure S5: Exclusion of single years Source: EU-SILC (authors’ calculations). Cohort-size coefficients are estimated as described in Section 3; the estimated model also includes region, year and age dummies; the solid line represents the cohort-size coefficients from the full model, the dashed lines the corresponding 95% confidence interval. -25 -20 -15 -10 -5 0 2012 2011 2010 2009 2008 2007 2006 2005 Unemployment share 0102030 40 Employment share 2012 2011 2010 2009 2008 2007 2006 2005 lower CI (95%) Coefficient upper CI (95%) lower CI (95%) Coefficient upper CI (95%) 143 Supplementary material Chapter 3 To further assess to what extent the estimated cohort-size effects vary between different groups of regions, we estimate Equation 3 separately for regions from three parts of Europe: Southern Europe (16 regions from Greece, Italy and Spain), Eastern Europe (14 regions from the Czech Republic, Estonia, Hungary, Latvia, Lithuania, Poland and Slovakia) and a combination of Northern and Western Europe (19 regions from Austria, Belgium, Denmark, France, Luxembourg and Sweden). Tables S1–S3 show the results of the baseline model as well as the coefficients from the model containing region-by-age dummies (with and without control variables). Estimating separate models for each of the three regions reduces the degrees of freedom compared to the pooled sample, which is reflected in higher standard errors. Moreover, the explanatory power of the instrument appears to be lower as evidenced by a reduction in the first-stage F-statistics. Nevertheless, in many specifications the 2SLS coefficients remain negative and significant in the unemployment model and positive and significant in the employment model when the Southern European regions are used. All of the coefficients have the expected sign and are significant at the 1% level for the sample of Eastern European regions. While there are no significant effects for the remaining regions of Northern and Western Europe, this need not imply that the relationship between cohort size and labour-market outcomes is structurally different in this part of Europe, but may rather be a reflection of the limited variation in the cohort-size variable as could already be seen in Figures A2 and A3. Figure S6: Exclusion of single age groups Source: EU-SILC (authors’ calculations). Cohort-size coefficients are estimated as described in Section 3; the estimated model also includes region, year and age dummies; the solid line represents the cohort-size coefficients from the full model, the dashed lines the corresponding 95% confidence interval. 010203040 29 28 27 26 25 Employment share -25 -20 -15 -10 -5 0 29 28 27 26 25 Unemployment share lower CI (95%) Coefficient upper CI (95%) lower CI (95%) Coefficient upper CI (95%) IAB-Bibliothek 367 144 Cohort size and youth labour-market outcomes: the role of measurement error Table S1: OLS and 2SLS regression results (Southern European regions) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -5.54 (3.58) -11.16** (5.13) -3.51 (3.68) -6.63 (5.36) -8.93** (3.86) -16.87*** (5.44) -6.56* (3.95) -12.23** (5.59) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 640 640 640 640 640 640 640 640 R20.53 0.52 0.55 0.55 0.57 0.57 0.59 0.59 F-stat – 338.11*** – 323.28*** – 340.42*** – 300.82*** ME(std) -0.02 -0.04** -0.01 -0.02 -0.03** -0.06*** -0.02* -0.04** Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 7.55* (4.05) 8.97 (6.58) 6.07 (4.11) 5.92 (7.27) 11.76** (4.61) 14.73** (6.78) 10.05** (4.69) 11.23 (7.24) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 640 640 640 640 640 640 640 640 R20.65 0.65 0.66 0.66 0.69 0.69 0.70 0.70 F-stat – 338.11*** – 323.28*** – 340.42*** – 300.82*** ME(std) 0.03* 0.03 0.02 0.02 0.04** 0.05** 0.03** 0.04 ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. 145 Supplementary material Chapter 3 Table S2: OLS and 2SLS regression results (Eastern European regions) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -9.21*** (1.87) -8.94*** (2.06) -5.78*** (2.00) -5.35*** (2.20) -9.90*** (2.18) -10.56*** (2.30) -5.96*** (2.29) -6.33*** (2.47) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 560 560 560 560 560 560 560 560 R20.32 0.32 0.40 0.40 0.36 0.36 0.44 0.44 F-stat – 843.57*** – 899.35*** – 705.96*** – 732.99*** ME(std) -0.02*** -0.02*** -0.01*** -0.01*** -0.02*** -0.02*** -0.01*** -0.01*** Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 15.91*** (2.33) 20.28*** (2.63) 11.99*** (2.49) 16.06*** (2.80) 14.74*** (2.64) 19.99*** (2.94) 9.87*** (2.84) 14.71*** (3.33) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 560 560 560 560 560 560 560 560 R20.41 0.41 0.48 0.48 0.46 0.45 0.53 0.52 F-stat – 843.57*** – 899.35*** – 705.96*** – 732.99*** ME(std) 0.03*** 0.04*** 0.02*** 0.03*** 0.03*** 0.04*** 0.02*** 0.03*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. IAB-Bibliothek 367 146 Cohort size and youth labour-market outcomes: the role of measurement error Table S3: OLS and 2SLS regression results (Northern and Western European regions) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -0.23 (3.93) 5.73 (7.88) 1.74 (4.07) 5.08 (7.51) 0.44 (4.22) 1.18 (7.51) 2.38 (4.46) 1.18 (6.99) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 759 759 759 759 759 759 759 759 R20.13 0.12 0.17 0.17 0.19 0.19 0.22 0.22 F-stat – 83.20*** – 94.78*** – 67.78*** – 77.30*** ME(std) -0.00 0.01 0.00 0.01 0.00 0.00 0.00 0.00 Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -1.55 (5.45) -0.23 (11.00) -2.74 (5.31) 4.75 (10.36) -6.66 (5.39) -3.28 (9.58) -7.75 (5.36) 1.20 (9.07) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 759 759 759 759 759 759 759 759 R20.24 0.24 0.30 0.29 0.32 0.32 0.37 0.37 F-stat – 83.20*** – 94.78*** – 67.78*** – 77.30*** ME(std) -0.00 -0.00 -0.01 0.01 -0.01 -0.01 -0.02 0.00 ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. 147 Supplementary material Chapter 3 S2.2 Robustness to changes in the model specification and the sample This part assesses the robustness of the estimated relationship between cohortsize and the unemployment and the employment shares to a variety of changes in the specification of the empirical model or the underlying sample. In Table S4 we first show that the paper’s results also hold when instead of aggregating the dependent variable to the level of the region-year-age group the underlying microdata is used (Angrist and Pischke, 2009). In this case the dependent variable is defined as a binary variable that indicates whether an individual i in age group j, region r and year t is unemployed (unempijrt ) or employed (empijrt ). In light of the strong assumptions that have to be made to ensure consistency in a binary dependent variable model with endogenous regressors (Cameron and Trivedi, 2009) and since the focus of the analysis is on estimating marginal effects rather than on making predictions, a linear probability model is used to which we apply the same IV estimation strategy that is outlined in Section 3. As the cohort-size variable is defined at a higher level of aggregation than the dependent variable, which now may also vary across individuals in the same region-year-age group, standard errors are clustered at the level of the region-age group cell (Moulton, 1990). Observations are weighted by the individual-level weights which have been provided as part of the EU-SILC data and which have then been calibrated so that the estimated size of a region-year-age-sex cell matches the population size as reported by Eurostat (see Section 2). The size of the standard errors increases compared to the aggregate-level analysis but all coefficients remain statistically significant at the 1% level. IAB-Bibliothek 367 148 Cohort size and youth labour-market outcomes: the role of measurement error Table S4: OLS and 2SLS regression results (individual-level analysis) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -10.32*** (1.97) -17.30*** (2.55) -8.38*** (1.95) -15.50*** (2.44) -12.84*** (2.35) -23.01*** (3.13) -10.17*** (2.34) -20.47*** (3.00) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations Individual-level Observations (cells) Region-year-age Region-age 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 R20.04 0.04 0.06 0.06 0.05 0.04 0.07 0.07 F-stat – 1,352.68*** – 1,568.95*** – 1,024.29*** - 1,385.52*** ME (std) -0.03*** -0.05*** -0.03*** -0.05*** -0.04*** -0.07*** -0.03*** -0.06*** Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 14.39*** (2.29) 24.32*** (3.05) 12.20*** (2.32) 22.06*** (3.02) 13.88*** (2.61) 26.55*** (3.71) 10.73*** (2.59) 23.15*** (3.63) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations Individual-level Observations (cells) Region-year-age Region-age 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 64,387 1,959 243 R20.07 0.07 0.10 0.10 0.08 0.08 0.10 0.10 F-stat – 1,352.68*** – 1,568.95*** – 1,024.29 – 1,385.52 ME(std) 0.04*** 0.08*** 0.04*** 0.07*** 0.04*** 0.08*** 0.03*** 0.07*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Standard errors that are clustered at the level of the region-age group cell are shown in parentheses. The regression is weighted using calibrated individual-level weights. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. 149 Supplementary material Chapter 3 In this paper a specific form of the cohort-size variable is used which, first, includes age groups that are up to two years younger and older and, second, assigns lower weights to age groups that are further away from the reference group. This specification is chosen to incorporate the assumption that members of an age group also compete with individuals that are slightly younger and older, but that substitutability decreases with the age difference. However, Wright (1991) already notes that this specific formulation is arbitrary. We therefore show that the results are robust to using a weighted cohort-size variable that only includes age groups that are up to one year younger or older (Equation S7), the relative size of the own-age group which does not consider any other age groups (Equation S8) as well as a three-year sum (Equation S9) and a five-year sum (Equation S10) in which each group receives an equal weight. Tables S5 to S8 show that the cohort-size coefficients retain their sign and significance. Since the distribution of these variables differ, it is useful to look at the marginal effects of a change in the corresponding cohort-size variable by one standard deviation instead of the cohort-size coefficients in order to compare the magnitude of the effects across the different specifications. CSjrt = (1/4)Nj – 1, rt + (1/2)Njrt + (1/4)Nj + 1 , rt N 16 – 64, rt [S7] CSjrt = Njrt N 16 – 64, rt [S8] CSjrt = Nj – 1, rt + Njrt + Nj + 1 , rt N16 – 64, rt [S9] CSjrt = Nj – 1, rt + Njrt + Nj + 1 , rt N16 – 64, rt [S10] IAB-Bibliothek 367 150 Cohort size and youth labour-market outcomes: the role of measurement error Table S5: OLS and 2SLS regression results (3-year weighted cohort-size variable) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -8.58*** (1.54) -16.30*** (2.00) -6.72*** (1.57) -14.20*** (1.94) -10.55*** (1.71) -21.97*** (2.29) -8.04*** (1.76) -18.98*** (2.16) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 1,959 1,959 1,959 1,959 1,959 1,959 1,959 1,959 R20.38 0.36 0.41 0.40 0.42 0.40 0.45 0.43 F-stat – 1,216.01*** – 1,242.94*** – 885.58*** – 972.80*** ME(std) -0.03*** -0.05*** -0.02*** -0.05*** -0.03*** -0.07*** -0.03*** -0.06*** Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 11.27*** (1.91) 23.02*** (2.51) 9.47*** (1.87) 21.04*** (2.39) 10.66*** (2.06) 25.23*** (2.78) 8.30*** (2.06) 22.40*** (2.59) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 1,959 1,959 1,959 1,959 1,959 1,959 1,959 1,959 R20.53 0.51 0.56 0.54 0.57 0.56 0.60 0.59 F-stat – 1,216.01*** – 1,242.94*** – 885.58*** – 972.80*** ME(std) 0.04*** 0.07*** 0.03*** 0.07*** 0.03*** 0.08*** 0.03*** 0.07*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation. 151 Supplementary material Chapter 3 Table S6: OLS and 2SLS regression results (own-age cohort-size variable) Panel A: Unemployment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size -3.07*** (0.95) -14.72*** (1.93) -1.99** (0.93) -12.96*** (1.88) -3.41*** (1.02) -20.02*** (2.32) -2.11*** (1.02) -17.62*** (2.18) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 1,959 1,959 1,959 1,959 1,959 1,959 1,959 1,959 R20.37 0.29 0.40 0.33 0.41 0.26 0.45 0.32 F-stat – 326.62*** – 337.13*** – 226.31*** – 249.71*** ME(std) -0.01*** -0.06*** -0.01** -0.05*** -0.01*** -0.08*** -0.01*** -0.07*** Panel B: Employment OLS 2SLS OLS 2SLS OLS 2SLS OLS 2SLS Cohort size 3.01** (1.23) 20.84*** (2.55) 1.96* (1.19) 19.29*** (2.41) 2.45* (1.26) 22.92*** (2.84) 1.21 (1.23) 20.73*** (2.64) Dummies Region Year Age Region-by-age Control variables Yes Yes Yes No No Yes Yes Yes No No Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes No Yes Yes Yes Yes No Yes Yes Yes Yes Yes Yes Yes Yes Yes Yes Observations (cells) Region-year-age 1,959 1,959 1,959 1,959 1,959 1,959 1,959 1,959 R20.52 0.42 0.55 0.46 0.57 0.46 0.60 0.50 F-stat – 326.62*** – 337.13*** – 226.31*** – 249.71*** ME(std) 0.01** 0.08*** 0.01* 0.08*** 0.01* 0.09*** 0.00 0.08*** ***/**/* indicate significance at the 1%/5%/10% level, respectively. Robust standard errors are shown in parentheses. The regression is weighted by the estimated number of male observations in a region-year-age cell. F-stat represents the first-stage F-statistic from a regression of the endogenous cohort-size variable on the instrument and control variables. ME(std) shows the change in the dependent variable if the cohort-size variable increases by one standard deviation.