scieee AI-readable full text Open interactive document viewer

The black box as a control for payoff-based learning in economic games

Burton-Chellew, Maxwell N.,West, Stuart A.

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Burton-Chellew, Maxwell N.; West, Stuart A. Article The black box as a control for payoff-based learning in economic games Games Provided in Cooperation with: MDPI – Multidisciplinary Digital Publishing Institute, Basel Suggested Citation: Burton-Chellew, Maxwell N.; West, Stuart A. (2022) : The black box as a control for payoff-based learning in economic games, Games, ISSN 2073-4336, MDPI, Basel, Vol. 13, Iss. 6, pp. 1-15, https://doi.org/10.3390/g13060076 This Version is available at: https://hdl.handle.net/10419/329987 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Citation: Burton-Chellew, M.N.; West, S.A. The Black Box as a Control for Payoff-Based Learning in Economic Games. Games 2022,13, 76. https://doi.org/10.3390/g13060076 Academic Editors: Kjell Hausken and Ulrich Berger Received: 29 September 2022 Accepted: 14 November 2022 Published: 16 November 2022 Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations. Copyright: © 2022 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https:// creativecommons.org/licenses/by/ 4.0/). games Article The Black Box as a Control for Payoff-Based Learning in Economic Games Maxwell N. Burton-Chellew 1,* and Stuart A. West 2 1Department of Economics, University of Lausanne, CH-1015 Lausanne, Switzerland 2Department of Biology, University of Oxford, Oxford OX1 3RB, UK *Correspondence: [email protected] Abstract: The black box method was developed as an “asocial control” to allow for payoff-based learning while eliminating social responses in repeated public goods games. Players are told they must decide how many virtual coins they want to input into a virtual black box that will provide uncertain returns. However, in truth, they are playing with each other in a repeated social game. By “black boxing” the game’s social aspects and payoff structure, the method creates a population of self-interested but ignorant or confused individuals that must learn the game’s payoffs. This lowinformation environment, stripped of social concerns, provides an alternative, empirically derived null hypothesis for testing social behaviours, as opposed to the theoretical predictions of rational self-interested agents (Homo economicus). However, a potential problem is that participants can unwittingly affect the learning of other participants. Here, we test a solution to this problem in a range of public goods games by making participants interact, unknowingly, with simulated players (“computerised black box”). We find no significant differences in rates of learning between the original and the computerised black box, therefore either method can be used to investigate learning in games. These results, along with the fact that simulated agents can be programmed to behave in different ways, mean that the computerised black box has great potential for complementing studies of how individuals and groups learn under different environments in social dilemmas. Keywords: altruism; asocial control; behavioural economics; conditional cooperation; confusion; directional learning; reinforcement learning; social preferences 1. Introduction Understanding human behaviour in social dilemmas is of crucial importance to solving many global issues [ 1 – 7 ]. Experiments using economic games provide a useful tool for investigating social behaviours [ 8 ]. By making participants pay for their decisions, experimenters hope to measure social preferences on the assumption that participants pay for preferred outcomes [ 9 ]. By using games with repeated decisions (repeated games), experimenters hope to measure how individuals respond to the behaviours of others (social responses) [ 10 – 18 ]. However, in repeated games, social responses can be confounded by individuals responding to their payoffs and learning how to play the game (payoffbased learning) [ 19 – 23 ]. This problem is particularly acute if many participants start the experiment without fully understanding the game’s payoffs [ 24 – 28 ]. Consequently, experimental control treatments are required to control for potentially confounding factors such as payoff-based learning. One solution to this problem of confounding is to make individuals face the same decision but in a low-information environment stripped of all social concerns (“asocial controls”) [ 19 , 25 , 27 , 29 – 33 ]. For example, Burton-Chellew et al. introduced the black box method as an “asocial control” to decouple social responses from payoff-based learning in repeated public-goods games [ 19 – 21 , 34 ]. Specifically, individuals interacted with a virtual black box, with which they could make voluntary inputs of “virtual coins” to Games 2022,13, 76. https://doi.org/10.3390/g13060076 https://www.mdpi.com/journal/games Games 2022,13, 76 2 of 15 obtain uncertain returns over several rounds 1 . However, in reality, the experiment actually involved groups of real participants playing a typical public goods game, with all the usual payoffs and social connections, just unknowingly 2 . By repeating the black box game for multiple rounds, one could measure how inputs evolved in populations of ignorant individuals with no social concerns [ 19 – 21 ]. The black box thus aimed to capture the psychology of self-interested but ignorant/confused individuals that use trial and error learning to improve their earnings. In this way, it was consistent with a rich history of prior studies that investigated how individuals learn in low-information environments and how reinforcement learning can affect cooperation [35–42]. The black box as an asocial control provided an alternative null hypothesis, empirically derived from behavioural observations, to the usual theoretical null hypotheses of a population of perfectly rational and selfish agents (Homo economicus). This “baseline” measure could then be compared to behaviour in versions of the normal, “revealed” public goods game to test if the addition of social information affected aggregate behaviour. BurtonChellew and West’s original results showed that aggregate contributions in the black box treatment were largely indistinguishable from those in the standard “revealed” public goods game, where individuals can observe their groupmates’ decisions, consistent with models of payoff-based learning [ 19 ]. In both cases, despite the income maximising decision in the one-shot version of the game being to contribute 0%, initial levels of contributions averaged around 40–50%, before gradually declining to approximately 15% by round 16, the final round. In both cases, most individuals contributed 0% in the final round, but approximately four percent of individuals still contributed fully. While these similarities did not confirm that individuals were using payoff-based learning, they did mean that one could not reject the null hypothesis of self-interest unless one assumed the participants perfectly understood the revealed game and that the similar levels of cooperation were mere coincidence. Although there were large similarities between the black box results and the typical results, it is important to keep in mind that the black box was providing a simplified model of behaviour based on the extreme assumption that all players are ignorant/confused and respond only to their own payoffs [ 19 ]. However, the black box also allows for examinations of how individuals learn and for estimating parameters within explicit hypothesised learning rules [ 20 , 21 ]. For example, subsequent collaborations with H. Nax and H. Peyton Young analysed individual-level data to estimate how much individuals value the earnings of their groupmates [ 20 ] and how individuals use payoff-based learning in the non-social and two social settings [ 21 ]. Not surprisingly, some differences in behaviour were found across the three treatments (the black box is, after all, like Homo economicus, a rather extreme hypothesis/model). When individuals could observe their groupmates’ decisions, they showed some conditional responses, but only if they could not also observe their groupmates’ payoffs (which is technically redundant information if individuals fully understand the game). However, payoff-based learning was significant in all three treatments [ 20 ], including even the manner of such learning [ 21 ], suggesting participants were motivated to try and increase their own income in all game forms. Together, these results suggested that conditional cooperation was more a function of social learning rather than a social preference for equal outcomes, although learning and social preferences may also interact [43]. The black box can also be easily modified and adapted to test different hypotheses. For example, Burton-Chellew and West [ 32 ] also subsequently used the black box to show that payoff-based learning is impeded when either group size (N) or the marginal per capita return from contributing is large (MPCR). By testing behaviour in three different black boxes that varied in either group size (N = 3 or 12) or the marginal per capita return from contributing (MPCR = 0.4 or 0.8), they showed that a large group size and/or a high MPCR and thus reduces the correlation between personal contributions and personal payoffs, thereby impeding payoff-based learning and potentially explaining why the rate of decline in cooperation varies across studies. They confirmed this hypothesis with a Games 2022,13, 76 3 of 15 comparative analysis that compared the rates of decline in 237 published public goods games. They found that rates of decline in contributions were slower when either group size or MPCR was large, and more specifically, when the estimated correlation between personal contributions and personal payoffs was weaker, a principle proved in their black box experiment [32]. However, one potential issue with the black box is that because participants are interacting, albeit unknowingly, the learning of one participant changes the learning environment for other participants. While this is also true for revealed social games, it may complicate efforts to discern individual learning from collective learning [ 44 ]. Another possible issue is that individuals do not know they can provide benefits to other participants, which may raise ethical concerns for some reviewers (however we do not think this omission of externalities constitutes deception). Here, we present a modified black box method that solves these two potential issues. Our solution is to make individuals still interact with a black box but change the setup so that individuals are grouped not with each other but with computerised players (computerised black box) (Figure 1). Otherwise, the set-up remains the same for the participants. There are several advantages to this approach: (1) individuals do not affect other participants, and thus can be treated as independent data points, providing more statistical power for given costs; (2) the learning environment can be maintained constant; (3) individuals are not affecting each other’s payoffs, thereby removing any potential ethical concerns; (4) computerized players receive no earnings making the study of behaviour in large groups more affordable; and (5) computerised players can be programmed to play in different, interesting ways, allowing one to test various hypotheses that would otherwise be unfeasible without using deception. We replicate the experimental design from Burton-Chellew and West, 2021, which used three different black boxes that varied in either group size or the cost of contributing to create one “easy” learning condition with a small group size and low MPCR (N = 3 and MPCR = 0.4) and two “difficult” learning conditions with either a large MPCR ( N=3 and MPCR = 0.8) or a large group size (N = 12 and MPCR = 0.4) [ 32 ]. Individuals could input 0–20 virtual coins in each round. However, instead of connecting human participants together, here we use computerised groupmates (programmed to input a random integer drawn from a uniform distribution of 0–20 coins). The payoff formula remained identical for all rounds and was the same for the human and computerised black boxes. This allowed us to compare rates of learning in the two methods, depending on both group size and MPCR. If behaviour with computerised black boxes qualitatively replicates behaviour with human black boxes, then the computerised black box method can be used as a complementary method to test hypotheses without any concerns about participants affecting each other’s behaviour and/or earnings. We also address a related research question on payoff-based learning. As mentioned above, Burton-Chellew and West, 2021, previously showed that payoff-based learning is impeded when groups are large or MPCR is high [ 32 ]. In such conditions, participants still contributed around 50% at the end of 16 rounds, despite the Nash equilibrium being 0%, indicative of zero learning. Nevertheless, it may be that individuals just need more time to learn in these challenging conditions and will eventually learn not to contribute. To test this, we repeated the black boxes with the difficult learning conditions, but under two conditions, a short game and a long game (16 versus 40 rounds). In all cases, we measured learning in two ways; (1) how quickly incentivised contributions converged towards the Nash equilibrium of 0 contributions, and (2) by asking participants at the end of the experiment to report their belief about what was the best number to contribute (“input”) into the black box (this was unincentivised). Games 2022,13, 76 4 of 15 Games 2022, 13, x FOR PEER REVIEW 3 of 17 decline in cooperation varies across studies. They confirmed this hypothesis with a comparative analysis that compared the rates of decline in 237 published public goods games. They found that rates of decline in contributions were slower when either group size or MPCR was large, and more specifically, when the estimated correlation between personal contributions and personal payoffs was weaker, a principle proved in their black box experiment [32]. However, one potential issue with the black box is that because participants are interacting, albeit unknowingly, the learning of one participant changes the learning environment for other participants. While this is also true for revealed social games, it may complicate efforts to discern individual learning from collective learning [44]. Another possible issue is that individuals do not know they can provide benefits to other participants, which may raise ethical concerns for some reviewers (however we do not think this omission of externalities constitutes deception). Here, we present a modified black box method that solves these two potential issues. Our solution is to make individuals still interact with a black box but change the set-up so that individuals are grouped not with each other but with computerised players (computerised black box) (Figure 1). Otherwise, the set-up remains the same for the participants. There are several advantages to this approach: (1) individuals do not affect other participants, and thus can be treated as independent data points, providing more statistical power for given costs; (2) the learning environment can be maintained constant; (3) individuals are not affecting each other’s payoffs, thereby removing any potential ethical concerns; (4) computerized players receive no earnings making the study of behaviour in large groups more affordable; and (5) computerised players can be programmed to play in different, interesting ways, allowing one to test various hypotheses that would otherwise be unfeasible without using deception. Figure 1. Black box methodologies. In the original black box, participants are connected online and interact in the usual experimental manner for economic games. However, by “black-boxing” the social aspect of the game or the game’s rules and payoffs, one can investigate how participants learn under certain conditions. If one is concerned about individuals affecting either the learning or the payoffs of other participants, one can replace the focal player’s interaction partners with Black box Input Output Black box Input Output Input Input Input Input Original black box Computerized black box Figure 1. Black box methodologies. In the original black box, participants are connected online and interact in the usual experimental manner for economic games. However, by “black-boxing” the social aspect of the game or the game’s rules and payoffs, one can investigate how participants learn under certain conditions. If one is concerned about individuals affecting either the learning or the payoffs of other participants, one can replace the focal player’s interaction partners with programmed computerised/virtual players (computerised black box). This also allows for more control over the learning environment, as partners can be programmed to be more/less cooperative, etc. 2. Results 2.1. Learning with Hidden Humans or Hidden Computers We found that there was no significant difference in contributions (“inputs”) depending upon whether individuals were grouped with humans or computers. The rate of decline in contributions across all 16 rounds of the short games did not significantly differ between the original black box with humans and the computerised black box in any of the three black boxes (Figure 2; Table 1). Specifically, the game round x groupmates interaction was non-significant in all three black boxes (generalised linear mixed models controlling for autocorrelation among groups/individuals: when N = 3 and MPCR = 0.4, Z = 0.2, p= 0.821 ; when N = 3 and MPCR = 0.8, Z = − 1.2, p= 0.222 ; and when N = 12 and MPCR = 0.4, Z = −0.3, p= 0.760, Table 1). As an additional check, we also compared final round contributions (“inputs”), which could be argued to be the best measure of learning. Again, we found no significant differences between playing with humans or with computers in any of the three black boxes (Table 2). Specifically, when learning was easy (N = 3 and MPCR = 0.4), mean ± SE final round inputs (0–20 virtual coins) were 3.0 ± 0.57 coins with humans and 4.7 ±0.90 coins with computers (Wilcoxon rank-sum test: W = 537, p= 0.797). When learning was difficult, mean final round inputs were typically around 50% (10 virtual coins) with both humans and computerised groupmates (N = 3 and MPCR = 0.8, with humans = 10.1 ± 0.91 coins, with computers = 9.4 ± 1.12 coins, W = 573.5, p= 0.563; N = 12 and MPCR = 0.4, with humans = 10.3 ±0.89 coins, with computers = 9.7 ±1.12 coins, W = 128, p= 0.806). Games 2022,13, 76 5 of 15 Games 2022, 13, x FOR PEER REVIEW 5 of 17 Figure 2. Learning in a black box, either with hidden human or computerised groupmates. We varied both the group size (N) and the benefit of contributing (MPCR) across three black boxes. Participants played all three black boxes, in counter-balanced order, but here we only show naïve behaviour (their first black box). Data show the mean contribution per round, with 95% confidence intervals based on the group means (in games with computers, the independent group is just one individual). The rate of learning was broadly similar regardless of playing with humans or computers. The linear regressions do not account for random effects/repeated measures and are therefore for illustration purposes only. The figures are annotated with the sample sizes of independent replicates (groups of humans or individuals grouped with computers). Table 1. Contributions over time. Analysis of how contributions (“inputs”) change during the game for each black box depending on if groupmates were humans or computers. Generalised linear mixed model with a binomial logit link, and random intercepts for both groups and individuals and random slopes for individuals. N = 3, MPCR = 0.4 N = 3, MPCR = 0.8 N = 12, MPCR = 0.4 Fixed Effects Z p Z p Z p Intercept (humans) 0.7 0.476 −1.4 0.159 3.6 <0.001 Round −8.0 <0.001 2.0 0.046 −1.0 0.298 Groupmates (computers) 0.3 0.795 2.2 0.030 −0.1 0.896 Round x Groupmates 0.2 0.821 −1.2 0.222 −0.3 0.760 N. obs. 1888 1856 1792 N. individuals 118 116 112 N. groups 70 68 46 Random effects Variance St. dev. Variance St. dev. Variance St. dev. Individual intercept 1.12 1.059 1.36 1.168 3.00 1.733 Individual slope 0.03 0.16 7 0.07 0.255 0.02 0.170 Group intercept 0.60 0.77 7 1.08 1.038 0.00 0.00 N = 24 N = 46 N = 24 N = 44 N = 6 N = 40 N =3, M PC R =0.4 N = 3 , M P C R = 0 .8 N=12, MPCR=0.4 116116116 0 25 50 75 100 Game round (1−16) Percent input Hidden groupmates Humans Computers Figure 2. Learning in a black box, either with hidden human or computerised groupmates. We varied both the group size (N) and the benefit of contributing (MPCR) across three black boxes. Participants played all three black boxes, in counter-balanced order, but here we only show naïve behaviour (their first black box). Data show the mean contribution per round, with 95% confidence intervals based on the group means (in games with computers, the independent group is just one individual). The rate of learning was broadly similar regardless of playing with humans or computers. The linear regressions do not account for random effects/repeated measures and are therefore for illustration purposes only. The figures are annotated with the sample sizes of independent replicates (groups of humans or individuals grouped with computers). Table 1. Contributions over time. Analysis of how contributions (“inputs”) change during the game for each black box depending on if groupmates were humans or computers. Generalised linear mixed model with a binomial logit link, and random intercepts for both groups and individuals and random slopes for individuals. N = 3, MPCR = 0.4 N = 3, MPCR = 0.8 N = 12, MPCR = 0.4 Fixed Effects Z pZpZp Intercept (humans) 0.7 0.476 −1.4 0.159 3.6 <0.001 Round −8.0 <0.001 2.0 0.046 −1.0 0.298 Groupmates (computers) 0.3 0.795 2.2 0.030 −0.1 0.896 Round x Groupmates 0.2 0.821 −1.2 0.222 −0.3 0.760 N. obs. 1888 1856 1792 N. individuals 118 116 112 N. groups 70 68 46 Random effects Variance St. dev. Variance St. dev. Variance St. dev. Individual intercept 1.12 1.059 1.36 1.168 3.00 1.733 Individual slope 0.03 0.167 0.07 0.255 0.02 0.170 Group intercept 0.60 0.777 1.08 1.038 0.00 0.00 Games 2022,13, 76 6 of 15 Table 2. Final contributions. Comparison of mean final round contributions (“inputs”) of virtual coins into the black box (0–20 coins). Comparisons made with Wilcoxon rank-sum test. Black Box Input: Humans Input: Computers (Short) Input: Computers (Long) W1P1W2P2 N = 3, MPCR = 0.4 3.0 ±0.57 4.7 ±0.90 / 573 0.797 / / N = 3, MPCR = 0.8 10.1 ±0.91 9.4 ±1.12 6.2 ±0.97 573.5 0.563 1287.5 0.025 N = 12, MPCR = 0.4 10.3 ±0.89 9.7 ±1.12 5.4 ±0.98 128 0.806 1238.5 0.006 1 Comparing final inputs with human or computerised groupmates. 2 Comparing final inputs in short or long games. We also asked the participants at the end of the 16 rounds if they thought there was a best number to input and if so, what it was (methods). Again, there were no significant differences depending on whether playing with humans or with computers (Figure 3; Table 3). Specifically, the mean ± SE stated beliefs (0–20 coins) for when N = 3 and MPCR = 0.4 were 1.5 ± 0.79 coins with humans and 3.0 ± 0.80 coins with computers (Wilcoxon rank-sum test, W = 348, p= 0.072); for when N = 3 and MPCR = 0.8, they were 8.8 ±1.32 coins with humans and 8.9 ± 1.59 coins with computers (W = 469.5, p= 0.888 ); and for when N = 12 and MPCR = 0.4, were 9.8 ± 1.38 coins with humans and 8.0 ±1.37 coins with computers (W = 387, p= 0.349). Games 2022, 13, x FOR PEER REVIEW 7 of 17 Figure 3. Groupmates and beliefs. Histograms show the frequency of each stated belief about what was the best number to input into the black box. Dashed vertical lines show the mean response. All responses are from naïve participants after finishing their first black box. The figures are annotated with the number of individuals. Table 3. Beliefs about the best number. The mean ±SE value participants stated as the best number to input at the end of the game. Comparisons made with Wilcoxon rank-sum test. Black Box Humans (N) * Computers—Short (N) Computers—Long (N) W 1 P 1 W 2 P 2 N = 3 , MPCR = 0.4 1.5 ± 0.79 (27) * 3.0 ± 0.80 (34) / 348 0.072 / / N = 3, MPCR = 0.8 8.8 ± 1.32 (40) 8.9 ± 1.59 (24) 3.5 ± 1.10 (24) 469.5 0.888 410.5 0.010 N = 12, MPCR = 0.4 9.8 ± 1.38 (28) * 8.0 ± 1.37 (24) 4.0 ± 0.93 (30) 387 0.349 496.5 0.016 * An error prevented data collection from some participants in the first seven sessions with human groupmates. 1 Comparing beliefs after play with human or computerised groupmates. 2 Comparing beliefs after short or long games. N = 27 N = 34 N = 40 N = 24 N = 28 N = 24 N=12, MPCR=0.4 N=3, M PCR=0.8 N=3, M PCR=0.4 0 2 4 6 8 10 12 14 16 18 20 0.0 0.2 0.4 0.6 0.0 0.2 0.4 0.6 0.0 0.2 0.4 0.6 Believed best input (0−20) Frequency Hidden groupmates Humans Computers Figure 3. Groupmates and beliefs. Histograms show the frequency of each stated belief about what was the best number to input into the black box. Dashed vertical lines show the mean response. All responses are from naïve participants after finishing their first black box. The figures are annotated with the number of individuals. Games 2022,13, 76 7 of 15 Table 3. Beliefs about the best number. The mean ± SE value participants stated as the best number to input at the end of the game. Comparisons made with Wilcoxon rank-sum test. Black Box Humans (N) * Computers— Short (N) Computers— Long (N) W1P1W2P2 N = 3, MPCR = 0.4 1.5 ±0.79 (27) * 3.0 ±0.80 (34) / 348 0.072 / / N = 3, MPCR = 0.8 8.8 ±1.32 (40) 8.9 ±1.59 (24) 3.5 ±1.10 (24) 469.5 0.888 410.5 0.010 N = 12, MPCR = 0.4 9.8 ±1.38 (28) * 8.0 ±1.37 (24) 4.0 ±0.93 (30) 387 0.349 496.5 0.016 * An error prevented data collection from some participants in the first seven sessions with human groupmates. 1 Comparing beliefs after play with human or computerised groupmates. 2 Comparing beliefs after short or long games. Overall, we found no significant differences between either inputs or beliefs, depending on if individuals were grouped with humans or computers. Our experiments with computerised groupmates replicated the results from the prior study with human groupmates [ 32 ]. Rates of learning were qualitatively similar regardless of groupmates being humans or computers in all three black box settings (Figure 2). These results mean that the original black box with humans can be used without having to worry too much about collective learning, or alternatively that the new, computerized, black box method can be used in certain contexts to obtain qualitatively similar results. However, we caution that for the “easy” black box (N = 3 and MPCR = 0.4), the final round inputs and the post-game beliefs about the value of the best input were lower, but not significantly, in the human black box. Looking at Figure 2, it may be that the learning rates for the “easy” black box ( N=3 and MPCR = 0.4) would have diverged if the experiment had continued for longer than 16 rounds, but we find no statistical support for this prediction within our data. 2.2. Learning in Longer Games We found clear evidence of payoff-based learning in the long-run games (Figure 4). Overall, the estimated rate of decline was significantly negative in both black boxes (Table 4, generalised linear mixed model controlling for individual: when N = 3 and MPCR = 0.8 , Z = −2.8 ,p= 0.005; when N = 12 and MPCR = 0.4, Z = − 3.3, p< 0.001, depending on black box). However, for both black boxes, the rate of decline was not significantly different between the short and long games, suggesting that the rate of learning is relatively constant within these time frames despite being undetectable in the short games (Table 4, round x game length interaction: N = 3 and MPCR = 0.8, Z = 1.0, p= 0.308; N = 12 and MPCR = 0.4, Z = 0.5, p= 0.643). Table 4. The effect of game length. Analysis of how inputs change during the game for each black box depending on game length (16 or 40 rounds). Generalised linear mixed model with a binomial logit link and random intercepts and slopes for individuals. N = 3, MPCR = 0.8 N = 12, MPCR = 0.4 Fixed Effects Z pZp Intercept 1.1 0.257 1.7 0.080 Round −2.8 0.005 −3.3 <0.001 Game length (short) 0.6 0.575 0.4 0.666 Round x Game length 1.0 0.308 0.5 0.643 N. obs. 2544 2480 N. individuals 90 86 N. groups 90 86 Random effects Variance St. dev. Variance St. dev. Individual intercept 3.00 1.731 5.67 2.381 Individual slope 0.01 0.116 0.01 0.122 Games 2022,13, 76 8 of 15 Games 2022, 13, x FOR PEER REVIEW 9 of 17 Figure 4. Learning and game length. The green data are the same as in Figure 2. Data show mean contributions with 95% confidence intervals, depending on game length, for two different black box parameter settings. The rate of learning was broadly similar in both black boxes regardless of game length, but because learning is slow in these parameter settings (large groups or high MPCR), the learning is only evident in long games. The linear regressions do not account for random effects/repeated measures and are therefore for illustration purposes only. The figures are annotated with the number of independent replicates (individuals grouped with computers). Table 4. The effect of game length. Analysis of how inputs change during the game for each black box depending on game length (16 or 40 rounds). Generalised linear mixed model with a binomial logit link and random intercepts and slopes for individuals. N = 3, MPCR = 0.8 N = 12, MPCR = 0.4 Fixed Effects Z p Z p Intercept 1.1 0.257 1.7 0.080 Round −2.8 0.005 −3.3 <0.001 Game length (short) 0.6 0.575 0.4 0.666 Round x Game length 1.0 0.308 0.5 0.643 N. obs. 2544 2480 N. individuals 90 86 N. groups 90 86 Random effects Variance St. dev. Variance St. dev. Individual intercept 3.00 1.731 5.67 2.381 N = 46 N = 44 N = 46 N = 40 N=12, MPCR=0.4 N=3, M PCR=0.8 0 8 16 24 32 40 0 25 50 75 100 0 25 50 75 100 Game round (1−40) Percent input Game length with computers Short Long Figure 4. Learning and game length. The green data are the same as in Figure 2. Data show mean contributions with 95% confidence intervals, depending on game length, for two different black box parameter settings. The rate of learning was broadly similar in both black boxes regardless of game length, but because learning is slow in these parameter settings (large groups or high MPCR), the learning is only evident in long games. The linear regressions do not account for random effects/repeated measures and are therefore for illustration purposes only. The figures are annotated with the number of independent replicates (individuals grouped with computers). Again, we compared the mean final contributions (“inputs”). These were significantly smaller, and thus closer to the income-maximising input of 0 coins, after the long game than after the short game in both black boxes (Table 2). Specifically, for the black box, where N = 3 and MPCR = 0.8, mean ± SE final inputs were 9.4 ± 1.12 coins in the short game and 6.2 ± 0.97 coins in the long game (Wilcoxon rank-sum test, W = 1287.5, p= 0.025). For the black box where N = 12 and MPCR = 0.4, final inputs were 9.7 ± 1.12 coins in the short game and 5.4 ±0.98 coins in the long game (W = 1238.5, p= 0.006). Moreover, when asked about their beliefs about a possible best number, stated beliefs were significantly lower on average at the end of the long games compared to the short games (Figure 5, Table 3). Specifically, the mean ± SE stated beliefs (0–20 coins) for when N=3 and MPCR = 0.8 , were 8.9 ± 1.59 coins in the short game and 3.5 ± 1.10 coins in the long game (Wilcoxon rank-sum test, W = 410.5, p= 0.010); for when N = 12 and MPCR = 0.4 , were 8.0 ± 1.37 coins in the short game and 4.0 ± 0.93 coins in the long game (W = 496.5, p= 0.016). Games 2022,13, 76 15 of 15 37. Zion, U.B.; Erev, I.; Haruvy, E.; Shavit, T. Adaptive behavior leads to under-diversification. J. Econ. Psychol. 2010 ,31, 985–995. [CrossRef] 38. Weber, R.A. ‘Learning’ with no feedback in a competitive guessing game. Games Econ. Behav. 2003,44, 134–144. [CrossRef] 39. Rapoport, A.; Seale, D.A.; Parco, J.E. Coordination in the Aggregate without Common Knowledge or Outcome Information. In Experimental Business Research; Zwick, R., Rapoport, A., Eds.; Springer: Boston, MA, USA, 2002; pp. 69–99. [CrossRef] 40. Colman, A.M.; Pulford, B.D.; Omtzigt, D.; Al-Nowaihi, A. Learning to cooperate without awareness in multiplayer minimal social situations. Cogn. Psychol. 2010,61, 201–227. [CrossRef] 41. Friedman, D.; Huck, S.; Oprea, R.; Weidenholzer, S. From imitation to collusion: Long-run learning in a low-information environment. J. Econ. Theory 2015,155, 185–205. [CrossRef] 42. Bereby-Meyer, Y.; Roth, A.E. The speed of learning in noisy games: Partial reinforcement and the sustainability of cooperation. Am. Econ. Rev. 2006,96, 1029–1042. [CrossRef] 43. Horita, Y.; Takezawa, M.; Inukai, K.; Kita, T.; Masuda, N. Reinforcement learning accounts for moody conditional cooperation behavior: Experimental results. Sci. Rep.-UK 2017,7, 39275. [CrossRef] 44. Peyton Young, H. Learning by trial and error. Games Econ. Behav. 2009,65, 626–643. [CrossRef] 45. Binmore, K. Why Experiment in Economics? Econ. J. 1999,109, F16–F24. [CrossRef] 46. Binmore, K. Economic man—Or straw man? Behav. Brain Sci. 2005,28, 817–818. [CrossRef] 47. Binmore, K. Why do people cooperate? Politics Philos. Econ. 2006,5, 81–96. [CrossRef] 48. Smith, V.L. Theory and experiment: What are the questions? J. Econ. Behav. Organ. 2010,73, 3–15. [CrossRef] 49. Friedman, D. Preferences, beliefs and equilibrium: What have experiments taught us? J. Econ. Behav. Organ. 2010 ,73, 29–33. [CrossRef] 50. Camerer, C.F. Experimental, cultural, and neural evidence of deliberate prosociality. Trends Cogn. Sci. 2013 ,17, 106–108. [CrossRef] 51. Fehr, E.; Schmidt, K.M. A theory of fairness, competition, and cooperation. Q. J. Econ. 1999,114, 817–868. [CrossRef] 52. Sobel, J. Interdependent preferences and reciprocity. J. Econ. Lit. 2005,43, 392–436. [CrossRef] 53. Saijo, T.; Nakamura, H. The Spite Dilemma in Voluntary Contribution Mechanism Experiments. J. Confl. Resolut. 1995 ,39, 535–560. [CrossRef] 54. Brunton, D.; Hasan, R.; Mestelman, S. The ‘spite’ dilemma: Spite or no spite, is there a dilemma? Econ. Lett. 2001 ,71, 405–412. [CrossRef] 55. Cherry, T.L.; Crocker, T.D.; Shogren, J.F. Rationality spillovers. J. Environ. Econ. Manag. 2003,45, 63–84. [CrossRef] 56. Fischbacher, U. z-Tree: Zurich toolbox for ready-made economic experiments. Exp. Econ. 2007,10, 171–178. [CrossRef] 57. Greiner, B. Subject pool recruitment procedures: Organizing experiments with ORSEE. J. Econ. Sci. Assoc. 2015 ,1, 114–125. [CrossRef] 58. Team, R. Integrated Development Environment for R; RStudio: Boston, MA, USA, 2020. 59. Burton-Chellew, M.N.; West, S.A. Data for: The black box as a control for payoff-based learning in economic games. Open Sci. Framew. 2022.