CS 3600: Markov Decision Processes in AI

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/28

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 9:48 PM on 8/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

29 Terms

1
New cards

Markov Decision Process (MDP)

Mathematical framework for decision-making in stochastic environments.

2
New cards

Grid World

A maze-like problem for testing MDPs.

<p>A maze-like problem for testing MDPs.</p>
3
New cards

Agent

Entity navigating the grid in Grid World.

4
New cards

States (S)

Different configurations or positions in the environment.

5
New cards

Actions (A)

Possible moves the agent can take.

6
New cards

Transitions (P)

Probability of moving from one state to another.

7
New cards

Rewards (R)

Feedback received after taking actions in states.

8
New cards

Discount Factor (γ)

Value representing future rewards' importance.

9
New cards

Policy (π)

Mapping from states to actions.

10
New cards

Utility

Sum of discounted rewards from a state.

11
New cards

Q-Values (Q*(s,a))

Expected utility from taking action a in state s.

12
New cards

Value Iteration

Method for computing optimal values iteratively.

<p>Method for computing optimal values iteratively.</p>
13
New cards

Policy Iteration

Method for improving policy through evaluation and extraction.

14
New cards

Bellman Equations

Recurrence relations defining optimal utility values.

15
New cards

Time-Limited Values (Vk)

Optimal value considering k remaining time steps.

<p>Optimal value considering k remaining time steps.</p>
16
New cards

Expectimax

Algorithm for decision-making under uncertainty.

17
New cards

Fixed Policy

A predetermined set of actions for states.

18
New cards

Utility under Fixed Policy (Vπ)

Expected total discounted rewards following a fixed policy.

19
New cards

Policy Evaluation

Calculating utilities for a fixed policy until convergence.

20
New cards

Policy Extraction

Deriving actions from optimal value estimates.

21
New cards

Convergence in Value Iteration

Vk values approach optimal values as iterations increase.

22
New cards

Optimal Policy (π*)

Best action to take from each state.

23
New cards

One-Step Lookahead

Evaluating immediate consequences of actions.

24
New cards

Car Racing MDP

Example illustrating MDP principles in a racing context.

25
New cards

Noisy Movement

Actions may not lead to intended outcomes.

26
New cards

Living Reward

Small reward received at each time step.

27
New cards

Big Rewards

Significant rewards received at the end of tasks.

28
New cards

Action Outcomes

Probabilities of resulting states from actions.

29
New cards

Learning vs. Planning

Learning involves acting to discover outcomes.