1/28
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Markov Decision Process (MDP)
Mathematical framework for decision-making in stochastic environments.
Grid World
A maze-like problem for testing MDPs.

Agent
Entity navigating the grid in Grid World.
States (S)
Different configurations or positions in the environment.
Actions (A)
Possible moves the agent can take.
Transitions (P)
Probability of moving from one state to another.
Rewards (R)
Feedback received after taking actions in states.
Discount Factor (γ)
Value representing future rewards' importance.
Policy (π)
Mapping from states to actions.
Utility
Sum of discounted rewards from a state.
Q-Values (Q*(s,a))
Expected utility from taking action a in state s.
Value Iteration
Method for computing optimal values iteratively.

Policy Iteration
Method for improving policy through evaluation and extraction.
Bellman Equations
Recurrence relations defining optimal utility values.
Time-Limited Values (Vk)
Optimal value considering k remaining time steps.

Expectimax
Algorithm for decision-making under uncertainty.
Fixed Policy
A predetermined set of actions for states.
Utility under Fixed Policy (Vπ)
Expected total discounted rewards following a fixed policy.
Policy Evaluation
Calculating utilities for a fixed policy until convergence.
Policy Extraction
Deriving actions from optimal value estimates.
Convergence in Value Iteration
Vk values approach optimal values as iterations increase.
Optimal Policy (π*)
Best action to take from each state.
One-Step Lookahead
Evaluating immediate consequences of actions.
Car Racing MDP
Example illustrating MDP principles in a racing context.
Noisy Movement
Actions may not lead to intended outcomes.
Living Reward
Small reward received at each time step.
Big Rewards
Significant rewards received at the end of tasks.
Action Outcomes
Probabilities of resulting states from actions.
Learning vs. Planning
Learning involves acting to discover outcomes.