Untitled Flashcards Set

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/8

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 4:57 PM on 12/15/24
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

9 Terms

1
New cards

Reinforcement Learning (RL)

Teaching a machine to learn by trial and error, using rewards (good) and penalties (bad) to improve decisions.

2
New cards

Q-Learning

An RL method where an agent learns a Q-table to decide the best actions in each situation.

3
New cards

Doing RL by Hand

Manually updating Q-values by picking actions (random/best), observing rewards, and repeating to improve decisions.

4
New cards

Bandits

A scenario where the agent chooses between options (like slot machines) to find the best reward.

5
New cards

Multi-Armed Bandits

Choosing between multiple options (e.g., slot machines), balancing exploration and exploitation.

6
New cards

Contextual Bandits

Bandits with context where the choice depends on the situation.

7
New cards

Epsilon-Greedy Strategy

A method to balance exploration and exploitation by exploring randomly with chance ε and exploiting the best-known action otherwise.

8
New cards

Upper Confidence Bound (UCB)

A strategy for balancing good rewards and uncertain actions.

9
New cards

Thompson Sampling

Uses probabilities to pick actions with the highest potential reward, balancing exploration and exploitation naturally.