Monday, May 20, 2013

What is to be done? Why reward is difficult to do without


The text below is the un-edited preprint of a commentary I wrote on Andy Clark's 'Whatever Next? Predictive Brains, Situated Agents, and the Future of Cognitive Science' in Behavioral and Brain Sciences. The published version of Clark's paper, with commentaries and a response, is here (probably behind a paywall.) Because behavioral and brain scientists were just falling over themselves to comment on Clark's target article, my commentary ended up appearing in Frontiers in Philosophical and Theoretical Psychology. That journal is open access, and so you can read the final version of my commentary at this link. You can read Clark's response to all of the commentaries at this link.

I've included the preprint text on this blog because it's one of the recent pieces in which I've directly argued for a thesis that has something to do with common currencies.

Commentary text


Clark’s synthesis of much recent work on sensory and motor systems in the brain is at once radical and curiously traditional. It is radical, among other things, concerning what representations are, how they are constructed, and what sensory and motor representations have in common. But it is traditionally cognitivist in viewing the main task of brains as being that of representing the world.

What this traditional orientation tends to neglect is the role of the brain as a system for selecting among available actions. This phenomenon has an ultimate aspect regarding the external standards relevant to assessing actions. Various behavioral ecological schemes for ranking actions in terms of their contribution to quantities such as fitness, and economic models of revealed preference, are the leading theoretical players here. The phenomenon also has a proximal aspect, which concerns the specific biological mechanisms, including neural ones, by means of which the values of different available actions might be represented, and selections between them made. On this topic the recent explosion of neuroeconomic research on decision processes in the brain is urgently relevant.

Natural agents have limited means of action, and those means have alternative – sometimes mutually exclusive – uses. That is to say the predicament of natural agents is fundamentally an economic one, even if it is not necessary that selection converge on a system for responding to the predicament in which economic variables are explicitly represented. Furthermore there is considerable evidence from behavioural ecology and other fields that many vertebrate behaviours in natural settings are economically efficient.

Neither the ultimate nor the proximal aspects of the problem of selecting between behaviours play a significant role in Clark’s account. Natural selection, fitness and biological descendents are not mentioned at all, and cognate concepts like adaptiveness feature in diluted form. There’s similarly little mention of decision and choice as theoretically understood in economics including neuroeconomics, none of incentives, and reward and utility appear only in the course of musing over whether it’s possible that cognitive neuroscience could do without reference to either (section 5.1). Clark does make some important points about action-centric representations, but even here does not consider the problem of action selection.

Of course, no survey can cover anything that anyone thinks is relevant, and it’s very easy to complain about things that are left out. Clark’s lack of engagement with neuroeconomics means missing a specific opportunity to make his general case even more compelling, because what is emerging in that field complements his case about sensory and motor systems in deep ways.

In his section (3.2) Clark apparently takes seriously the concern that an agent with the sort of brain that he’s been describing would be expected to ‘seek a nice dark room and stay in it’. Clark disposes of the worry by pointing out that creatures with real biological needs should ‘expect’ to follow exploratory strategies, and that these expectations themselves should recruit both perception and action. This is part of a reasonable and interesting response, but action selection under those conditions (as with most others) would still require some way of dealing with specific questions, such as where and how to forage, and how to trade off foraging with other expected behaviours such as predator avoidance and reproduction.

A related move appears later, in section (5.1) when he considers an austere vision of cognition that does without reference to goals and rewards, in favour of comprehensive analysis in terms of expectations. Clark correctly holds back from endorsing this possibility, but for relatively generic reasons to the effect that even if some description is in principle replaceable, it may be convenient to continue using it. This misses the main chance. Recent work on the neural implementation of decision in various vertebrates including humans has produced a body of results highly congenial to the unifying vision Clark supports.

Consider saccadic movements in rhesus monkeys. A key component in the neural implementation of these movements is the lateral intraparietal area (LIP), which comprises a topographic map integrating locations in the visual field and aspects of the muscular plans that would effect the centering of gaze on those locations. It, along with a network of other maps with varying topographies in the frontal eye fields, superior colliculus and related areas, provides a striking illustration of what Clark calls an ‘action-centric’ representation. In addition, as studies including Platt & Glimcher (1999) and Dorris & Glimcher (2004) have shown, some activity in LIP neurons of rhesus monkeys on visually identical trials varies in precise ways with the relative expected rewards (or relative subjective value) from saccades to the represented location. These representations are not merely ‘action-centric’ insofar as they combine answers to the questions ‘where is it?’ with ‘how do I gaze at it?’ They also include identifiable activity corresponding to the answer to ‘what’s it probably worth for me to look at it?’

There’s more. The expected relative reward values attached to saccadic and other movements are not sui generis. They’re predictions, and ones that get updated in the light of ongoing experience. Among the key findings on this topic is that dopamine neurons do not – as previously supposed – directly encode hedonic value (because if they did they would respond in the same to expected and unexpected rewards of equivalent hedonic worth). Rather it turns out that they encode some aspects of the difference between experienced and expected reward (Montague et al 1997, see also Bayer & Glimcher 2005). While many details about the operation of this system, and its interaction with other neural systems, have yet to be determined, it is nonetheless clear that crucial features of the neural systems for attaching values to sensory events and actions operate by means of prediction error. In this respect they suggest a way of expanding the scope of Clark’s claim about the importance of minimizing prediction error as a general goal of neural systems.

References

Bayer, H.M. and Glimcher, P.W. (2005). Midbrain dopamine neurons encode a quantitative reward prediction error signal. Neuron 47, 1 – 13.

Dorris, M.C. and Glimcher, P.W. (2004). Activity in posterior parietal cortex is correlated with the subjective desirability of an action. Neuron 44, 365 – 378.

Montague , P.R., Dayan , P., and Sejnowski, T.J. (1997). A framework for mesencephalic dopamine systems based on predictive Hebbian learning. J. Neurosci. 16, 1936 – 1947.

Platt, M.L. and Glimcher, P.W. (1999). Neural correlates of decision variables in parietal cortex. Nature 400, 233 – 238.

Version notes:
First posted May 20, 2013. Link to Clark's reply to commentaries added May 21, 2013.

Things that get suggested as common currencies


One of the things that makes the common currency topic interesting, is the variety of candidate currencies that get suggested.

Here is a list, that I’ll add to and flesh out periodically, of things that have been proposed as currencies:

Utility


In one of the economists’ technical senses, utility is not a psychological notion, but refers to whatever an agents behaviour tends to make more probable, as revealed in the pattern of behaviour itself.

An ‘agent’ on this view is anything whose activity reveals a consistent set of preferences, although additional assumptions can be introduced so that apparently inconsistent entities still turn out to be agents.

A consistent set of preferences revealed by an agent, represents the common currency of that agent. (Any two options will stand in some relation of relative preferability.)

There are other technical notions of utility (for example from behavioural economics) that I’ll list here in the future.

Reward/reinforcement


This is reinforcement in the behavioural psychologists' sense, where a reinforcer is something that changes the rate at which a behaviour is emitted. (In proper terminology it is behaviours that are reinforced, while agents are rewarded.)

So the various reinforcement values of all things that are reinforcing will also define a common currency, representing the effectiveness of all the incentives to which an agent is responsive.

An exemplary statement of the argument for a common currency in this setting is the following:
"In natural settings, the goals competing for behavior are complex, multidimensional objects and outcomes. Yet, for orderly choice to be possible, the utility of all competing resources must be represented on a single, common dimension" (Shizgal & Conover 1996).
There's a discussion of Shizgal and Conover's paper here.


Pain and Pleasure

I have not read this book.

According to Bentham pain and pleasure were the “two sovereign masters” which both explained human action, and enabled good and bad consequences of actions to be identified. Contemporary economics and behavioural psychology make considerably less reference to pleasure and pain than a Benthamite might have expected. Instead they focus on largely behavioural notions of utility and reward. Among those whose primary interest is pleasure and pain, one still finds suggestions that these constitute a ‘common currency’:
“Consistent with the idea that a common currency of emotion enables the comparison of pain and pleasure in the brain, the evidence reviewed here points to there being extensive overlap in the neural circuitry and chemistry of pain and pleasure processing at the systems level.” (Leknes & Tracey, 2008, p314).

Fitness


Fitness is (roughly and informally) relative propensity to have descendants. It is relative to competing variants in a population, and so typically defined by reference to genotypes.

If we focus specifically on the contribution to fitness of the behaviour of an individual, then we might go on to think that the contribution of individual behaviours (or behavioural dispositions) might be ordered with respect to how much the add or, or undermine, total fitness. Something along these lines is one goals of at least some behavioural ecologists:

“Any attempt to understand behavior in terms of the evolutionary advantage that it might confer has to find a "common currency" for comparing the costs and benefits of various alternative courses of action” (McNamara and Houston 1986: 358).

ATP


ATP is the standard abbreviation for Adenosine triphosphate. It’s the molecule used to transport energy in cells. It might seem like an odd thing to have on this list, although it is common to hear it referred to as the ‘energy currency’ of intracellular processes. Since any behavior has some energy cost (or gain), and any allocation of metabolic resources has an opportunity cost (in actions rendered unavailable, or made possible) it’s plausible to think that in a highly aggregated way, ATP represents an important overall budgetary bottom line.

Quite how to relate this to some of the other proposed currencies is another matter.


References


Leknes, S. and Tracey, I. (2008) A common neurobiology for pain and pleasure, Nature reviews: Neuroscience, 9,  pp314-320. [Publisher link.] [Google scholar citations.]

McNamara, J.M. and Houston, A.I. (1986) The Common Currency for Behavioral Decisions, The American Naturalist, 127(3), pp358-378. [Publisher link.] [Google scholar citations.]

Shizgal, P. & Conover, K. 1998. On the neural computation of utility, Current Directions in Psychological Science, 5(2), pp. 37-43. [Publisher’s site –may be behind a paywall] [Preprint version at CogPrints] [Google Scholar Citations]

Thursday, May 16, 2013

On the Motivational Commensurability of Pleasure and Pain


'On the Motivational Commensurability of Pleasure and Pain' is the working title of a talk/paper that is part of this sprawling ‘common currency’ project.

When the project started, I wasn’t especially interested in pleasure and pain. (That is, I wasn’t especially academically interested in them.) That was because I partly shared the behaviourist tendency in both the psychological and economic sense: introspection and affect weren’t much scientific use, and what we needed to know about motivation could be revealed in behaviour, or perhaps in behaviour and cognitive neuroscience.

One thing that changed my mind about this was finding out that pleasure and pain themselves get mooted as being the ‘common currency’. So, as a comparative investigator of common currencies, I had to pay some attention.

Another thing that changed my mind was paying a bit more attention to the pre-behaviourist history of some of the disciplines that are relevant to the project. For early utilitarians (like Bentham, and Mill) what was motivating coincided with what was pleasant or painful. Behaviourist psychology and most economics (since around the 1930s) have studiously avoided reference to affect (in favour of reinforcement, reward, and a de-psychologised notion of ‘utility’). But it’s an interesting question how what is discovered about motivation and what is discovered about pleasure and pain relate.

This represents what I somewhat pretentiously called (in version 1 of the talk) the ‘Optimistic Correspondence Hypothesis’, and which is fairly clearly false some of the time. This is my own diagram, but you will find equivalent images in some theoretically-minded discussions of pain.


Finally, I saw a call for papers for a future special issue of the Review of Philosophy and Psychology that got me thinking. I don’t know if I’ll have something I’m happy with ready in time for the deadline, but I started collecting some thoughts, and put together a working talk.

Presentations

15 May 2013, at UKZN philosophy. This was the first outing. I think it went pretty well, in the sense that I was reasonably clear, and (as I often find) immediately after thought of a number of improvements to the presentation and ways of describing the issues. So I’ll update the thing, and look for another audience I can pester with it. I'll post the slides here after I've refined them a little.


Monday, May 6, 2013

Cui bono? Selfish goals need to pay their way


The text below is the un-edited preprint of a commentary I've just submitted to BBS (Behavioral and Brain Sciences). The commentary is about a target article by Julie Y. Huang and John A. Bargh called "The Selfish Goal: Autonomously Operating Motivational Structures as the Proximate Cause of Human Judgment and Behavior".

I have no idea when the commentary might see print. It's a standard joke to say that 'BBS' stands for 'Badly Behind Schedule', and my commentary on Andy Clark's forthcoming target article (see 'My Own Efforts') appeared in another journal before the article it was commenting on. (There was a great flood of people wanting to say something, and some of those who didn't make the first cut had a second change through Frontiers in Theoretical and Philosophical and Psychology.)

As you'll see, I think that Huang and Bargh's proposal faces some serious difficulties. And some of them relate to common currencies.

Abstract

The target article risks not explaining the phenomena, including motivational conflict, that it claims to. The two mains reasons for this are: (1) it is unclear in what sense goals are‘selfish’. (2) We need an account of how selfish goals motivate people. If selfish goals aren’t in the replication business, then what is in it for them? And if they don’t offer people something that they want, how do they ever influence what people do?


Main Text

The proposal in the target article risks not explaining the phenomena, including motivational conflict, that it claims to. The two mains reasons for this are: (1) it is not clear enough in what sense goals might be ‘selfish’, or what incentives they respond to. (2) We need an account of how selfish goals motivate individual people, including how they compete with other incentives to which people respond.

Goals are apparently selfish in a way ‘analogous’ to Dawkins’ (1976) use with reference to genes. Dawkins argued that to understand much of biology it was necessary to take the perspective of the units of heredity. These are ‘selfish’ in the sense that their only interest is in being replicated. Serving this interest sometimes makes demands that oppose the well being of the individual carrying them. Dawkins focuses mostly on units of biological inheritance – genes – but also speculates that there may be analogously selfish cultural replicators or ‘memes’. In the case of memes, as with genes, selection processes favour efficient replicators, irrespective of whether this serves the interest of the ‘host’. In both cases it is clear in what sense the replicators are selfish, and what is ‘in it for them’. The game of life pays replicators with copies of themselves, and the various individuals in which they occur are means to that end. The interests of these replicators are not appreciated by the replicators themselves, but well defined in evolutionary game theory where payoffs are numbers of descendants (Maynard-Smith 1982).

The target article carefully avoids using the term ‘meme’, and makes almost no reference to replication (except in the context of glossing Dawkins’ popularisation of gene-centric selection). This suggests that the sense in which goals are selfish analogously with genes is not the same as Dawkins’ existing meme analogy. But if selfish goals don’t have a primary interest in replication, what are their interests? Goals, we are told, represent ‘end states’, and succeed by being pursued and/or by being completed.

This proposal needs to answer two crucial questions. First, what is ‘in it’ for goals such that pursuit and completion constitutes an incentive? Second, how might selfish goals compete for control of an individual?

It is difficult to discern an answer to the first question, or even clues as to what it might be, in the target article. If selfish goals aren’t replicators and are furthermore in some sense separate from the person they occupy, then they can neither be rewarded with copies of themselves, nor in whatever subjective utility the person responds to. (In any event, in the latter case they would not be selfish at all – they would simply be the person’s preferences.)

The second question is, if anything, even more important. The various (finite) degrees of freedom of any human represent a scarce resource that has alternative uses. That is to say the problem of behaviour allocation is essentially economic (Shizgal, 2012). Some of the processes that implement allocation are peripheral and relatively encapsulated, but most are not. The behaviour allocations of people are undoubtedly sensitive to costs and payoffs, even if it is a matter of controversy what specific economic model humans instantiate. In addition, there is mounting evidence that contemplating or selecting both desirable and aversive options in a wide range of modalities (including money, delayed money, risky money, food, drink, pain, looking at attractive faces, and social reputation) is consistently associated with activity proportional to behaviourally inferred desirability in a single brain region (for a recent review see Levy & Glimcher 2012). Independent of this specific evidence, the motor areas of the brain constitute a final common path for control processes, plausibly requiring any candidate deployment of the agent’s capacities to compete on the same terms as the others.

These considerations apply to selfish goals. The options available to a person (including end-states of goals) are often composed of complex mixtures of components (money, food, sex, status…) and in different modalities. Their availability could be immediate or delayed, and they can be subject to risk. The costs of options also vary in magnitude and type (effort, money, pain, delay…), and may themselves be multi-modal (e.g. including both monetary cost and delay). Making choices in even an approximately efficient way requires trading off these multi-modal options by reference to their net costs and benefits. Arguably, solving that problem is a significant part of what brains are for. Unless selfish goals are to be epiphenomenal, therefore, they need to compete along with the already recognized sources of subjective utility (or reward, or reinforcement), and they need to do so somewhere along the recognised pathways for behavioural control. 

These factors appear to make almost no impact on the argument in the target article, where there is no mention whatsoever of utility, incentive, consumption, pleasure or risk, and merely solitary and passing references to pain, reinforcement, and reward. But selfish goals need to motivate the agents that they occupy, and to do so they need to pay their way in some kind of incentive to which those agents are responsive.

Finally, it is worth pointing out that even non-selfish goals can come into conflict, and that people motivated purely by their own incentives can exhibit inconsistency over time. The most banal of objectives can be mutually exclusive (resting and working, specialising and generalising). All that is needed to explain some conflict, that is, is that not all desires (even those of a self) can be satisfied. In addition, the ways we price in the costs of various ways of being separated from a reward (by delay, risk, effort, etc.) can themselves be a source of inconsistency. This has been most extensively studied in the case of delayed rewards, where there is considerable evidence that humans discount rewards that will be received later (Ainslie 1992, Kable & Glimcher 2007) in a manner that corresponds more closely to a hyperbolic than exponential function. The mere passage of time, that is, can change the ranking of preferences within a single person.

References

Ainslie, G. (1992) Picoeconomics: The strategic interaction of successive motivational states within the person. Cambridge University Press.

Kable , J. W. , and Glimcher , P. W. ( 2007). The neural correlates of subjective value during intertemporal choice . Nature Neuroscience, 10(12): 1625-1633.

Levy, D. J. and Glimcher, P. W. (2012) The root of all value: a neural common currency for choice, Current Opinion in Neurobiology, 22: 1027-1038.

Maynard Smith, J. (1982) Evolution and the Theory of Games. Cambridge University Press.

Shizgal P (2012) Scarce means with alternative uses: Robbins’ definition of economics and its extension to the behavioral and neurobiological study of animal decision making. Frontiers in Neuroscience. 6(20). 


Wednesday, May 1, 2013

Shizgal and Conover’s ‘Orderly choice’ argument and evidence


ResearchBlogging.org One of the clearest statements of an argument for a common currency thesis is found early in Peter Shizgal and Kent Conover’s 1996 paper ‘On the neural computation of utility.’ This is a really important paper for the common currency topic, and should be required reading for anyone in the area. It doesn’t only state a version of one of the key arguments, it also provides an exemplary kind of empirical evidence.

Here is how they state the argument:

In natural settings, the goals competing for behavior are complex, multidimensional objects and outcomes. Yet, for orderly choice to be possible, the utility of all competing resources must be represented on a single, common dimension (Shizgal & Conover 1996).

The second sentence can be slightly reformulated (making one implicit premise explicit) as follows:

Premise 1: Orderly choice is possible.
Premise 2: For orderly choice to be possible, the utility of all competing resources must be represented on a single, common dimension.
So
Conclusion: The utility of all competing resources is represented on a single, common dimension.

This certainly looks like a valid argument. But what does it mean?

Well, the suggestion is clearly that (a) the rewards available to an organism might vary in lots of ways, across multiple dimensions (consider sex, rest, food of different kinds, grooming, water, avoiding predators, feeding young, etc.), and (b) making ‘orderly’ decisions between available rewards, the rewards all need to have a simple one-dimensional representation. (See The (very) Basic Big Idea on this site.)

Shizgal and Conover don’t claim that all choice is in fact orderly. But they clearly intend to say that when it is orderly, a common dimension of comparison is required. But what is it for choice to be orderly? Looking at Shizgal and Conover’s experiments will help clarify this.

The research Shizgal and Conover report on concerns ‘brain self-reward’. A famous paper by Olds and Milner (1954) reported that rats would work, including learning novel behaviours, for no more reinforcement than pulses of electrical stimulation to a part of the medial forebrain where an electrode terminated. Some early reports of brain-self reward – as it came to be called, or (BSR) for short – focused on ways in which it was unusual. The popular imagination got excited by the rat that pressed its lever nearly continually for around 20 days at about one press every two seconds (Valenstein & Beer 1964), or cases where food was apparently ignored over brain self-reward (Routtenberg & Lindy 1965). Shizgal and Conover sought, in part, to demonstrate that brain self-reward is a reinforcer like any other. In a series of experiments they had rats chose between trains of BSR pulses and infusions of sucrose solution.

The experimental design infused the juice directly into the rats’ mouths to give the sucrose solution a key property of BSR, which is that a single action leads to both procurement and consumption. In addition the swallowed fruit juice was drained away, reducing postingestive effects so that the infusions shared another property of BSR, which is absence of satiation.

Schematic illustration of experiment 1 from preprint version of the paper.

Shizgal and Conover manipulated the strength of reinforcement to each modality (by changing the duration or number of pulses in a train of BSR, or the size of the sucrose infusion) and measured the relative allocation of lever presses to each. Among other things they found that each individual reward modality was preferred over nothing, more of it preferred to less, and that increasing the opportunity cost of either (by increasing the reward available from the other lever) led to less consumption of that reward.

In a variation on the experiment they manipulated the magnitude of a BSR-only reward, while the alternative reward was a constant combination of BSR and sucrose solution. In this condition it took more BSR to make the rat forgo the compound reward than had been necessary for the sucrose part of the reward alone.

Schematic illustration of experiment 2 from preprint version of the paper.
Their conclusion is that the results of the first experiment imply, “that on a given trial, the rat selected the alternative that registered a larger value in a common system of measurement”, rather than following a categorical rule (such as ‘whenever the size of one reward is above some threshold, choose it’).

Regarding the second experiment with the combination rewards, they say it implies “that the electrical stimulation and the sucrose were subjected to a common evaluation”.

(The paper describes two further experiments, complementary to those I’ve just described. I’ll discuss them in a future posting.)

So, one way of glossing this is that Shizgal and Conover demonstrate that rat choices between BSR and sucrose solution is orderly in two senses:

  • Choices between modalities are quantitatively sensitive to opportunity cost.
  • Choices between rewards in one modality and two-modality combination rewards are similarly sensitive to the combined opportunity cost.

These results are important. They show (in terminology that I’ve described elsewhere on this site, but which isn’t used by Shizgal and Conover) that at least some rat choices have a certain patterning, conforming to an ultimate common currency.

Shizgal and Conover also clearly intend the conclusion that the results about an ultimate common currency support the view that there is a proximal one:
“We speculate that this common ability arises from a common action of the gustatory and electrical stimuli on a neural system that determines goal selection by signaling the utility of competing goals.”
The other two experiments, and their ongoing research programme, provide further support for this speculation. But that’s all I’ve got time for now.

References

Olds, J. & Milner, P. 1954. Positive reinforcement produced by electrical stimulation of septal area and other regions of rat brain. Journal of Comparative and Physiological Psychology, 47, pp. 419-427.
Routtenberg, A. & Lindy, J. 1965. Effects of the availability of rewarding septal and hypothalamic stimulation on bar pressing for food under conditions of deprivation. Journal of Comparative and Physiological Psychology, 60, pp. 158-161.
Shizgal, P. & Conover, K. 1998. On the neural computation of utility, Current Directions in Psychological Science, 5(2), pp. 37-43. [Publisher’s site –may be behind a paywall] [Preprint version at CogPrints] [Google Scholar Citations]
Valenstein, E.S. & Beer, B. 1964. Continuous opportunity for reinforcing brain stimulation. Journal for the Experimental Analysis of Behavior, 7, pp. 183-184.

Forthcoming attractions

Discussion of Shizgal (1999).

See also

Currencies can be 'ultimate' or 'proximal'




Shizgal, P., & Conover, K. (1996). On the Neural Computation of Utility. Current Directions in Psychological Science, 5 (2), 37-43 DOI: 10.1111/1467-8721.ep10772715

Friday, April 26, 2013

The invisible blog

So, the term 'common currency' is mostly used to talk about currency unions. This is, of course, where more than one political entity - usually a state - use the same currency in the sense of money.

The term used much less frequently in the sense relevant to this blog, which is with reference to decision, pattern in choices, and the processes by which choices are produced. As a result this blog is, so far, barely visible. Except for social traffic from people who pay attention to my personal activity on Facebook and Twitter, there's almost no traffic at all. Part of the problem is the dominance of what for me are irrelevant uses of 'common currency'. Even searching for "common currency brain science blog" doesn't (yet) return results where this site is anywhere near the top.

I'm going to try to take a few steps to improve matters. One of those is signing the blog up on Research Blogging, so that appropriate posts can be digested and appear in their feeds. I'll also see what I can do to generate some more inbound links.