RLHF and Prompt Engineering
Reinforcement Learning from Human Feedback
- RLHF uses signals from human evaluators (e.g., upvotes/downvotes) to improve AI model performance.
- Feedback helps the model adjust decisions to align with human expectations.
- Useful when defining a clear objective function is difficult.
- Example: social media moderation AI uses RLHF to better identify offensive or harmful content and limit its spread.
Prompt Engineering and the New Interaction Model
- Prompt engineering is a valuable tool with broad applications; it's a new interaction model.
- It changes how we talk to AI and shape its behavior.
- Goal: learn how to talk to AI.