NeurEcon Week 6 - Algorithms, RPE

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/29

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 11:35 PM on 10/9/25
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

30 Terms

1
New cards

what are latent variables

reward and prediction are unobservable, rewards and predictions are latent variables we cannot observe directly. comes with assumptions like there is a linear assocaition of reward to stimulus and assume predictions are driven by learning algorithm

2
New cards

what is the central idea behind reward prediction error RPE?

Dopamine neurons don’t just respond to reward itself —
they respond to the difference between what you expected and what you actually got. This difference is called the reward prediction error

Dopamine activity is positively related to experienced reward
and negatively related to predicted reward > when you experience a reward, dopamine neurons increase, when you predicted a reward or already excepted a reward dopamine neurons decrease 


Situation

What happens

Dopamine response

Better than expected

Reward is larger or earlier than expected

Dopamine burst (positive prediction error)

As expected

Reward matches what was predicted

No change in dopamine (zero error)

Worse than expected

Reward is smaller, later, or absent

Dopamine dip (negative prediction error)


3
New cards
<p>what was the goal of bush and mosteller?</p>

what was the goal of bush and mosteller?

mathamatically represent data representing how a rats response changes over conditioning stages. acquisition models: aim to capture how the probability of a conditioned response increases with pairings; introduce trial-by-trial association updates

as more and more cue-reward pairings go on, the prob of rat making association increases, can we model this?

Alast trial= association stregnth of most recent trial

Anext trial = assocaitoin strength of upcoming trial

alph a= learning rate 0>r<1

r = reward strength

x-axis = trials

y-axis = associative strength

Alast trial = 0 if you are on trial 1


weakness: no inclusion of time

<p>mathamatically represent data representing how a rats response changes over conditioning stages. acquisition models: aim to capture how the probability of a conditioned response increases with pairings; introduce trial-by-trial association updates</p><p>as more and more cue-reward pairings go on, the prob of rat making association increases, can we model this?</p><p>Alast trial= association stregnth of most recent trial</p><p>Anext trial = assocaitoin strength of upcoming trial</p><p>alph a= learning rate 0&gt;r&lt;1</p><p>r = reward strength</p><p>x-axis =&nbsp;trials</p><p>y-axis = associative strength</p><p>Alast trial = 0 if you are on trial 1</p><p></p><p>weakness: no inclusion of time</p>
4
New cards

what do you notice about the assocaitive strength as trials go on with Bush and Mosteller model?


what impact does alpha have on learning?

assocaitive strength converges to expected value EV of reward. model predicts shape of acquisitoin curve

—

alpha <1 allows system to remember the past, and to compute EV you need to track past trials

alpha =1 learning occurs too fast and model takes on an osillatory graph with asscicated strength bouncing between two values

<p>assocaitive strength converges to expected value EV of reward. model predicts shape of acquisitoin curve</p><p>—</p><p>alpha &lt;1 allows system to remember the past, and to compute EV you need to track past trials</p><p>alpha =1 learning occurs too fast and model takes on an osillatory graph with asscicated strength bouncing between two values</p>
5
New cards

what is the main idea behind temporal difference in temporal difference learning TDRL model? Sutton and BARTO (1998)

the difference between reward recieved and reward expected.

V(st)new = assocaitive strength new trial

V(st)old =value estimate for currnent stat

V(st+1) = current state # + 1

t = time

alpha = learning rate

gamma (y) = discounting parameter, 1

r = reward delivered in current state

<p>the difference between reward recieved and reward expected.</p><p>V(st)new = assocaitive strength new trial</p><p>V(st)old =value estimate for currnent stat</p><p>V(st+1) = current state # + 1</p><p>t = time</p><p>alpha = learning rate</p><p>gamma (y) = discounting parameter, 1</p><p>r = reward delivered in current state</p>
6
New cards
<p>what part of the TDLR eqaution accounts for reward prediction error?</p><p>what does the presence of the cue do for learning?</p><p>does TDLR account for dopamine diminished respond to reward as it becomes expected?</p>

what part of the TDLR eqaution accounts for reward prediction error?

what does the presence of the cue do for learning?

does TDLR account for dopamine diminished respond to reward as it becomes expected?

RPE = [ r+yV(st+1) - V(stcurrent) ]

r+yV(st+1) > information about actual current outcome and future expectatoin

V(stcurrent) > current expectation

—

presence of the cue is a prediction error signal

—

yes!

on trial 1 state 10, PE = reward. over the trials, PE continues to decline at state 10 as prediction error is reduced.

<p>RPE = [ r+yV(st+1) - V(stcurrent) ]</p><p>r+yV(st+1) &gt; information about actual current outcome and future expectatoin</p><p>V(stcurrent) &gt; current expectation</p><p>—</p><p>presence of the cue is a prediction error signal</p><p>—</p><p>yes!</p><p>on trial 1 state 10, PE = reward. over the trials, PE continues to decline at state 10 as prediction error is reduced.</p>
7
New cards

Schultz, Dayan & Montague (Science, 1997) Prediction Error representation

knowt flashcard image


This is the foundational study that connected dopamine neuron firing to the mathematical reward prediction error from the TDRL (Temporal difference reinforcement leanring) algorithm.


3D plot shows how prediction error changes across

time (x-axis)

learning trials (y-axis)


early on dopamine burst happen at reward (end of trial) but over more trials bursts happen early in trial when cue persented

<p>This is the foundational study that connected <strong>dopamine neuron firing</strong> to the <strong>mathematical reward prediction error</strong> from the <strong>TDRL (Temporal difference reinforcement leanring) algorithm</strong>.</p><p></p><p>3D plot shows how prediction error changes across</p><p>time (x-axis)</p><p>learning trials (y-axis)</p><p></p><p>early on dopamine burst happen at reward (end of trial) but over more trials bursts happen early in trial when cue persented</p>
8
New cards
term image
  • Early trials: spike at reward (unexpected reward → big positive PE).

  • Later trials: spike shifts earlier (cue predicts reward → PE at cue).

  • Reward omitted: dopamine dips below baseline (negative PE = worse than expected).

  • Positive PE: unexpected reward → dopamine burst

  • Zero PE: predicted reward → no dopamine change

  • Negative PE: missing reward → dopamine dip


9
New cards

the temporal difference reinforcement learning TDRL alg shows us how to build a representation of value with prediction errors. it predicts there will be prediction error signals at the time of — and —- under the right conditions.

do dopamiine neurons confirm to the PE signal predicted by algorithm?

what kind of framework positive or normative is TDRL?

it predicts there will be prediction error signals at the time of cues and rewards under the right condition.


yes dopamine neurons confirm to the PE signal predicted by algorithm


TDRL is a normative framework of how to solve reinforcement learning probems. normative fraemwork describe how a system should work, if something is a TDRL system, it should behave in specific way > we can use the framework to test if biological brain deviates from a TDRL framework

10
New cards

what are the 3 levels of analysis David Marr and Tomas Peggio suggest?

computational level - what does system do and why


algorithm level - how does system achieve computaitonal goal


implementaltion level - how does physical or biological sytem suport algorithm



<p>computational level - what does system do and why</p><p></p><p>algorithm level - how does system achieve computaitonal goal</p><p></p><p>implementaltion level - how does physical or biological sytem suport algorithm</p><p></p><p></p>
11
New cards

what does the TDRL use to update values estimates?

updates value estimates based on prediction error signals, this algorithm is physically implemented in dopamine neurons.

12
New cards

Caplin et Dean (2008) - Axiomatic methods, dopamine, and reward prediction error


what is the economic approach to defining reward prediction error RPE?

why is aciomatic approach needed?

axiom1 : coherent prize dominance

axiom 2: coherent lottery dominance

axiom 3: no suprise equivalence


—

prediction and reward and RPE are latent variables (they are infered), we need to relatelatend variables to something we can observe, axioms specify formal properties that the RPE must satisfy

13
New cards

axiom 1 : coherent prize dominance

axiom1 : coherent prize dominance; for a fized lottery better prizes should always lead to higher dopamine release no matter the probability


a lottery with 10 and 20 prize, x-axis = probability/chance of winning prize 1

<p>axiom1 : coherent prize dominance; for a fized lottery better prizes should always lead to higher dopamine release no matter the probability</p><p></p><p>a lottery with 10 and 20 prize, x-axis = probability/chance of winning&nbsp;prize 1</p>
14
New cards

axiom 2: coherent lottery dominance

for a fixed prize, lotteries having a higher probability of returning the better prize should always lead to lower dopamine release

as prob prize 1 increase to the right, dopmaine response observed when prize 1 or prize 2 is observed should decrease


because dopmaine does not decrease across x-axis for both prizes, it violates axioms

<p>for a fixed prize, lotteries having a higher probability of returning the better prize should always lead to lower dopamine release</p><p>as prob prize 1 increase to the right, dopmaine response observed when prize 1&nbsp;or prize 2 is observed should decrease </p><p></p><p>because dopmaine does not decrease across x-axis for both prizes, it violates axioms</p>
15
New cards

axiom 3: no suprise equivalence

If an outcome is fully predictable (no surprise), the brain’s dopamine response should be the same no matter how big or small the reward is — because there’s no prediction error to signal.

dopamine response should be same for both prizes if they are 100% expected

but in graph, we see that 100 chance of winning prize 1 and 0 percent chance of winning prize 2 does not have same level of activity.

<p>If an outcome is <strong>fully predictable</strong> (no surprise), the brain’s <strong>dopamine response</strong> should be <strong>the same</strong> no matter how big or small the reward is — because there’s <strong>no prediction error</strong> to signal.</p><p>dopamine response should be same for both prizes if they are 100% expected</p><p>but in graph, we see that 100 chance of winning prize 1 and 0 percent chance of winning prize 2 does not have same level of activity.</p>
16
New cards
<p>this graph is what should look like if dopeamine response follows all 3 axioms, how so?</p>

this graph is what should look like if dopeamine response follows all 3 axioms, how so?

100 prob of winning prize 1 and 100 prob winning prize 2 have same dopamine release (baseline/no change since its not a suprise) (axiom 3)


having a higher prob of lottery 1 (better prize) leads to lower dopamine release (axiom 2)


when both prizes fully expected, better prize (1) should always lead to higher dopamine releas (axiom 3)

17
New cards

what was Caplin et Dean (2008) behaivor task - Axiomatic methods, dopamine, and reward prediction error

  • gave subjects $100

  • On each trial, they face 50% chance to win or lose $5 (a simple gamble).

  • The task focuses on brain activity during the outcome phase (when participants find out if they won or lost).

  • Purpose: see whether nucleus accumbens activity behaves like an RPE signal (consistent with the 3 axioms)


  • When the outcome is better than expected (win $5):

    • dopamine-related regions (nucleus accumbens) should increase activity → positive RPE.

  • When the outcome is worse than expected (lose $5):

    • activity should decrease → negative RPE.

  • When outcomes are fully predictable, activity should flatten → zero RPE.


<ul><li><p>gave subjects $100</p></li></ul><ul><li><p>On each trial, they face <strong>50% chance to win or lose $5</strong> (a simple gamble).</p></li><li><p>The task focuses on brain activity during the <strong>outcome phase</strong> (when participants find out if they won or lost).</p></li><li><p>Purpose: see whether <strong>nucleus accumbens activity</strong> behaves like an RPE signal (consistent with the 3 axioms)</p></li></ul><p></p><ul><li><p>When the outcome is <strong>better than expected</strong> (win $5):</p><ul><li><p>dopamine-related regions (nucleus accumbens) should <strong>increase activity</strong> → positive RPE.</p></li></ul></li><li><p>When the outcome is <strong>worse than expected</strong> (lose $5):</p><ul><li><p>activity should <strong>decrease</strong> → negative RPE.</p></li></ul></li><li><p>When outcomes are <strong>fully predictable</strong>, activity should flatten → zero RPE.</p></li></ul><p></p>
18
New cards
<p><strong>Caplin et Dean (2008) - Axiomatic methods, dopamine, and reward prediction error </strong></p><p></p><p><strong>signal in FMRI data of NAcc</strong></p>

Caplin et Dean (2008) - Axiomatic methods, dopamine, and reward prediction error


signal in FMRI data of NAcc

activity in Nacc with FMRI > signal correlates with axioms, dopmaine concentration changes correlate with axioms

during outcome, red lines show smaller probablity gives bigge response cause we were less likely to win but won. blue lines show negative RPE, outcome worse tha expected acitivty dips (gagged)

19
New cards

Does this mean that NAcc are generation prediction error signals?

Hart et al 2014 used a fast scan cylic voltammetry to measure dopamine concentration in real time, fMRI signal suggest that input to signal from dopamine neurons is the result of NAcc response, conection through dendrites makes it os AP is proportion to RPE signal generated by dopamine neurons in VTA but NAcc is not source of RPE. VTA spikes lead to dopamine changes in downstream targets. 

<p>Hart et al 2014 used a fast scan cylic voltammetry to measure dopamine concentration in real time, fMRI signal suggest that input to signal from dopamine neurons is the result of NAcc response, conection through dendrites makes it os AP is proportion to RPE signal generated by dopamine neurons in VTA but NAcc is not source of RPE. VTA spikes lead to dopamine changes in downstream targets.&nbsp;</p>
20
New cards
<p>what does this represent</p>

what does this represent

associative value after conditioning/ learning

21
New cards

what was question and experiment proposed by Maes et al 2020

wants to see if after light as assocaitive value, if it can behave as a reward, can prediction signal elicited by light cause learning of new cues that precede light?


a new neutral cue (tone) presented before the light, no juice delivered, if dopamine fires during the light it has a teaching signal, the tone should gain value.

22
New cards

What’s the behavioral evidence for SOC?

Why is SOC important for studying dopamine?

experiment: can pairing a novel neutral stimulus (tone) with a conditioned stimulus (light) cause conditioned responding


After training, the animal responds to the new cue (tone) alone, showing that it learned an association even though no reward was given during tone–light pairings.




rats develop conditioned resposne (CS) to the tone alone. when dopamine signal supressed at the time of light/tone due no learning occurs, tells us that the signal at the time of the cue
is required for SOC learning.


Two camps interpret the signal differently
1) It encodes a prediction error signal?
2) It encodes a cues associative
importance/value?

23
New cards

What are the two main hypotheses about the cue-evoked dopamine signal?

  • Prediction Error (PE) hypothesis: dopamine firing reflects a teaching signal that updates learning.

  • Value (V) hypothesis: dopamine firing reflects the stored value of the cue (how good it is), not a learning signal itself.


24
New cards

How is optogenetics and blocking used in the SOC experiment?

two groups, one no opto and one opto suppression

no opto group: normal dopamine signal when light appears  > animal should learn tone > light > reward and show conditioned response to tone

opto suppression: dopa neuron supressed when light presented during tone-light pairing

When dopamine neurons were suppressed at the light, the animal did not learn the second order tone association — showing that the cue-evoked dopamine burst is necessary for new learning (SOC). it is a teaching signal not just value signal.

25
New cards

How can we tell if the cue signal represents “prediction error” vs. “value”?

By using blocking experiments. If the cue signal reflects prediction error, blocking that signal prevents learning; if it reflects value, blocking it should actually increase surprise later and allow RPE at reward so learning.

26
New cards
<p>describe blocking expreiment</p>

describe blocking expreiment

conditioning: cue A (light) > juice

blocking: new cue (Y) added . AY > juice, AX > juice, BZ > juice

  • opto supression during AX cue, optosupression during reward after AY.

test: AY, AX, BZ

the green triangle represents optogenetic supression, artificual dopamine stimulation

<p>conditioning: cue A (light) &gt; juice</p><p> </p><p>blocking: new cue (Y) added . AY &gt; juice, AX &gt; juice, BZ &gt; juice</p><ul><li><p>opto supression during AX cue, optosupression during reward after AY.</p></li></ul><p>test: AY, AX, BZ</p><p>the green triangle represents optogenetic supression, artificual dopamine stimulation</p>
27
New cards
<p>if the signal is reward value for prediction, then we prevent V from forming</p>

if the signal is reward value for prediction, then we prevent V from forming

If we think the signal is reward value for prediction
then we prevent V from forming.
Since no V is present, when the reward is delivered, we get
a positive PE at the time of reward (RPE).
This leads to the tone entering into an association and we get
learning

some say value signal is at the time of the cue, dope neurons doing TDLR agorithm, value signals are needed for learning but the cut itself is not the value signal. 

<p><span>If we think the signal is reward value for prediction</span><span><br></span><span>then we prevent V from forming.</span><span><br></span><span>Since no V is present, when the reward is delivered, we get</span><span><br></span><span>a positive PE at the time of reward (RPE).</span><span><br></span><span>This leads to the tone entering into an association and we get</span><span><br></span><span>learning</span><span><br></span></p><p>some say value signal is at the time of the cue, dope neurons doing TDLR agorithm, value signals are needed for learning but the cut itself is not the value signal.&nbsp;</p>
28
New cards
<p>if the signal/ cue is a prediction error, then the RPE portion of the alrogithm is deleted. why?</p>

if the signal/ cue is a prediction error, then the RPE portion of the alrogithm is deleted. why?

if the signal is PE, we prevent computaition of PE at the time of cue but we don’t impact value (assocaitbie strength) Since V is present, when the reward is delivered, we get no RPE at the time of reward > the tone is thus blocked from forming an association so no learning


we are inhbiting pomutation of PE not representation of V and

<p>if the signal is PE, we prevent computaition of PE at the time of cue but we don’t impact value (assocaitbie strength) Since V is present, when the reward is delivered, we get no RPE at the time of reward &gt; the tone is thus blocked from forming an association so no learning</p><p></p><p>we are inhbiting pomutation of PE not representation of V and</p>
29
New cards
term image

This experiment shows that dopamine at the cue is not just a “stored value” signal — it’s a teaching signal (prediction error).
When you silence dopamine neurons at the cue, the brain fails to learn new associations because it can’t compute the “surprise” that drives learning.

30
New cards

if signal at cue is a value signal then what happens to value representation and PE representation at downstream states?


if signal at cue is a RPE signal then what happens to value representation and PE representation at downstream states?

RPE @ cue > value is 4 from moment 4-10, PE = 0 from moemnt 4-10


value @ cue > the value signal is 0 until we get to moment 10, RPE signal is 0 until moment 4, value is 0 from moment 4-10.