1/29
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
what are latent variables
reward and prediction are unobservable, rewards and predictions are latent variables we cannot observe directly. comes with assumptions like there is a linear assocaition of reward to stimulus and assume predictions are driven by learning algorithm
what is the central idea behind reward prediction error RPE?
Dopamine neurons don’t just respond to reward itself —
they respond to the difference between what you expected and what you actually got. This difference is called the reward prediction error
Dopamine activity is positively related to experienced reward
and negatively related to predicted reward > when you experience a reward, dopamine neurons increase, when you predicted a reward or already excepted a reward dopamine neurons decreaseÂ
Situation | What happens | Dopamine response |
|---|---|---|
Better than expected | Reward is larger or earlier than expected | Dopamine burst (positive prediction error) |
As expected | Reward matches what was predicted | No change in dopamine (zero error) |
Worse than expected | Reward is smaller, later, or absent | Dopamine dip (negative prediction error) |

what was the goal of bush and mosteller?
mathamatically represent data representing how a rats response changes over conditioning stages. acquisition models: aim to capture how the probability of a conditioned response increases with pairings; introduce trial-by-trial association updates
as more and more cue-reward pairings go on, the prob of rat making association increases, can we model this?
Alast trial= association stregnth of most recent trial
Anext trial = assocaitoin strength of upcoming trial
alph a= learning rate 0>r<1
r = reward strength
x-axis =Â trials
y-axis = associative strength
Alast trial = 0 if you are on trial 1
weakness: no inclusion of time

what do you notice about the assocaitive strength as trials go on with Bush and Mosteller model?
what impact does alpha have on learning?
assocaitive strength converges to expected value EV of reward. model predicts shape of acquisitoin curve
—
alpha <1 allows system to remember the past, and to compute EV you need to track past trials
alpha =1 learning occurs too fast and model takes on an osillatory graph with asscicated strength bouncing between two values

what is the main idea behind temporal difference in temporal difference learning TDRL model? Sutton and BARTO (1998)
the difference between reward recieved and reward expected.
V(st)new = assocaitive strength new trial
V(st)old =value estimate for currnent stat
V(st+1) = current state # + 1
t = time
alpha = learning rate
gamma (y) = discounting parameter, 1
r = reward delivered in current state


what part of the TDLR eqaution accounts for reward prediction error?
what does the presence of the cue do for learning?
does TDLR account for dopamine diminished respond to reward as it becomes expected?
RPE = [ r+yV(st+1) - V(stcurrent) ]
r+yV(st+1) > information about actual current outcome and future expectatoin
V(stcurrent) > current expectation
—
presence of the cue is a prediction error signal
—
yes!
on trial 1 state 10, PE = reward. over the trials, PE continues to decline at state 10 as prediction error is reduced.
![<p>RPE = [ r+yV(st+1) - V(stcurrent) ]</p><p>r+yV(st+1) > information about actual current outcome and future expectatoin</p><p>V(stcurrent) > current expectation</p><p>—</p><p>presence of the cue is a prediction error signal</p><p>—</p><p>yes!</p><p>on trial 1 state 10, PE = reward. over the trials, PE continues to decline at state 10 as prediction error is reduced.</p>](https://knowt-user-attachments.s3.amazonaws.com/659b62d3-bf00-4b69-af00-9088030e4d10.png)
Schultz, Dayan & Montague (Science, 1997) Prediction Error representation

This is the foundational study that connected dopamine neuron firing to the mathematical reward prediction error from the TDRL (Temporal difference reinforcement leanring) algorithm.
3D plot shows how prediction error changes across
time (x-axis)
learning trials (y-axis)
early on dopamine burst happen at reward (end of trial) but over more trials bursts happen early in trial when cue persented


Early trials: spike at reward (unexpected reward → big positive PE).
Later trials: spike shifts earlier (cue predicts reward → PE at cue).
Reward omitted: dopamine dips below baseline (negative PE = worse than expected).
Positive PE: unexpected reward → dopamine burst
Zero PE: predicted reward → no dopamine change
Negative PE: missing reward → dopamine dip
the temporal difference reinforcement learning TDRL alg shows us how to build a representation of value with prediction errors. it predicts there will be prediction error signals at the time of — and —- under the right conditions.
do dopamiine neurons confirm to the PE signal predicted by algorithm?
what kind of framework positive or normative is TDRL?
it predicts there will be prediction error signals at the time of cues and rewards under the right condition.
yes dopamine neurons confirm to the PE signal predicted by algorithm
TDRL is a normative framework of how to solve reinforcement learning probems. normative fraemwork describe how a system should work, if something is a TDRL system, it should behave in specific way > we can use the framework to test if biological brain deviates from a TDRL framework
what are the 3 levels of analysis David Marr and Tomas Peggio suggest?
computational level - what does system do and why
algorithm level - how does system achieve computaitonal goal
implementaltion level - how does physical or biological sytem suport algorithm

what does the TDRL use to update values estimates?
updates value estimates based on prediction error signals, this algorithm is physically implemented in dopamine neurons.
Caplin et Dean (2008) - Axiomatic methods, dopamine, and reward prediction error
what is the economic approach to defining reward prediction error RPE?
why is aciomatic approach needed?
axiom1 : coherent prize dominance
axiom 2: coherent lottery dominance
axiom 3: no suprise equivalence
—
prediction and reward and RPE are latent variables (they are infered), we need to relatelatend variables to something we can observe, axioms specify formal properties that the RPE must satisfy
axiom 1 : coherent prize dominance
axiom1 : coherent prize dominance; for a fized lottery better prizes should always lead to higher dopamine release no matter the probability
a lottery with 10 and 20 prize, x-axis = probability/chance of winning prize 1

axiom 2: coherent lottery dominance
for a fixed prize, lotteries having a higher probability of returning the better prize should always lead to lower dopamine release
as prob prize 1 increase to the right, dopmaine response observed when prize 1Â or prize 2 is observed should decrease
because dopmaine does not decrease across x-axis for both prizes, it violates axioms

axiom 3: no suprise equivalence
If an outcome is fully predictable (no surprise), the brain’s dopamine response should be the same no matter how big or small the reward is — because there’s no prediction error to signal.
dopamine response should be same for both prizes if they are 100% expected
but in graph, we see that 100 chance of winning prize 1 and 0 percent chance of winning prize 2 does not have same level of activity.


this graph is what should look like if dopeamine response follows all 3 axioms, how so?
100 prob of winning prize 1 and 100 prob winning prize 2 have same dopamine release (baseline/no change since its not a suprise) (axiom 3)
having a higher prob of lottery 1 (better prize) leads to lower dopamine release (axiom 2)
when both prizes fully expected, better prize (1) should always lead to higher dopamine releas (axiom 3)
what was Caplin et Dean (2008) behaivor task - Axiomatic methods, dopamine, and reward prediction error
gave subjects $100
On each trial, they face 50% chance to win or lose $5 (a simple gamble).
The task focuses on brain activity during the outcome phase (when participants find out if they won or lost).
Purpose: see whether nucleus accumbens activity behaves like an RPE signal (consistent with the 3 axioms)
When the outcome is better than expected (win $5):
dopamine-related regions (nucleus accumbens) should increase activity → positive RPE.
When the outcome is worse than expected (lose $5):
activity should decrease → negative RPE.
When outcomes are fully predictable, activity should flatten → zero RPE.


Caplin et Dean (2008) - Axiomatic methods, dopamine, and reward prediction error
signal in FMRI data of NAcc
activity in Nacc with FMRI > signal correlates with axioms, dopmaine concentration changes correlate with axioms
during outcome, red lines show smaller probablity gives bigge response cause we were less likely to win but won. blue lines show negative RPE, outcome worse tha expected acitivty dips (gagged)
Does this mean that NAcc are generation prediction error signals?
Hart et al 2014 used a fast scan cylic voltammetry to measure dopamine concentration in real time, fMRI signal suggest that input to signal from dopamine neurons is the result of NAcc response, conection through dendrites makes it os AP is proportion to RPE signal generated by dopamine neurons in VTA but NAcc is not source of RPE. VTA spikes lead to dopamine changes in downstream targets.Â


what does this represent
associative value after conditioning/ learning
what was question and experiment proposed by Maes et al 2020
wants to see if after light as assocaitive value, if it can behave as a reward, can prediction signal elicited by light cause learning of new cues that precede light?
a new neutral cue (tone) presented before the light, no juice delivered, if dopamine fires during the light it has a teaching signal, the tone should gain value.
What’s the behavioral evidence for SOC?
Why is SOC important for studying dopamine?
experiment:Â can pairing a novel neutral stimulus (tone) with a conditioned stimulus (light) cause conditioned responding
After training, the animal responds to the new cue (tone) alone, showing that it learned an association even though no reward was given during tone–light pairings.
rats develop conditioned resposne (CS) to the tone alone. when dopamine signal supressed at the time of light/tone due no learning occurs, tells us that the signal at the time of the cue
is required for SOC learning.
Two camps interpret the signal differently
1) It encodes a prediction error signal?
2) It encodes a cues associative
importance/value?
What are the two main hypotheses about the cue-evoked dopamine signal?
Prediction Error (PE) hypothesis: dopamine firing reflects a teaching signal that updates learning.
Value (V) hypothesis: dopamine firing reflects the stored value of the cue (how good it is), not a learning signal itself.
How is optogenetics and blocking used in the SOC experiment?
two groups, one no opto and one opto suppression
no opto group: normal dopamine signal when light appears > animal should learn tone > light > reward and show conditioned response to tone
opto suppression: dopa neuron supressed when light presented during tone-light pairing
When dopamine neurons were suppressed at the light, the animal did not learn the second order tone association — showing that the cue-evoked dopamine burst is necessary for new learning (SOC). it is a teaching signal not just value signal.
How can we tell if the cue signal represents “prediction error” vs. “value”?
By using blocking experiments. If the cue signal reflects prediction error, blocking that signal prevents learning; if it reflects value, blocking it should actually increase surprise later and allow RPE at reward so learning.

describe blocking expreiment
conditioning: cue A (light) > juice
blocking: new cue (Y) added . AY > juice, AX > juice, BZ > juice
opto supression during AX cue, optosupression during reward after AY.
test: AY, AX, BZ
the green triangle represents optogenetic supression, artificual dopamine stimulation


if the signal is reward value for prediction, then we prevent V from forming
If we think the signal is reward value for prediction
then we prevent V from forming.
Since no V is present, when the reward is delivered, we get
a positive PE at the time of reward (RPE).
This leads to the tone entering into an association and we get
learning
some say value signal is at the time of the cue, dope neurons doing TDLR agorithm, value signals are needed for learning but the cut itself is not the value signal.Â


if the signal/ cue is a prediction error, then the RPE portion of the alrogithm is deleted. why?
if the signal is PE, we prevent computaition of PE at the time of cue but we don’t impact value (assocaitbie strength) Since V is present, when the reward is delivered, we get no RPE at the time of reward > the tone is thus blocked from forming an association so no learning
we are inhbiting pomutation of PE not representation of V and


This experiment shows that dopamine at the cue is not just a “stored value” signal — it’s a teaching signal (prediction error).
When you silence dopamine neurons at the cue, the brain fails to learn new associations because it can’t compute the “surprise” that drives learning.
if signal at cue is a value signal then what happens to value representation and PE representation at downstream states?
if signal at cue is a RPE signal then what happens to value representation and PE representation at downstream states?
RPE @ cue > value is 4 from moment 4-10, PE = 0 from moemnt 4-10
value @ cue > the value signal is 0 until we get to moment 10, RPE signal is 0 until moment 4, value is 0 from moment 4-10.