11. Safety Measures

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/7

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:47 PM on 9/29/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

8 Terms

1
New cards
What are the 9 AI safety concerns and their main mitigation actions?

SC1 Data distribution ≠ real world → MA1 well-justified data acquisition.
SC2 Distribution shift over time → MA10 continuous learning/updating.
SC3 Incomprehensible behavior → MA3 gray-box/explainability.
SC4 Unknown behavior in rare critical situations → MA5 structured testing + MA6 deep analysis of test results.
SC5 Unreliable confidence → MA2 reliable/calibrated confidence.
SC6 Brittleness of DNNs → MA4 threat modelling and defenses.
SC7 Inadequate train/test separation → MA7 data partitioning guidelines.
SC8 Dependence on labeling quality → MA8 labeling guidelines.
SC9 Safety not considered in metrics → MA9 safety-aware evaluation metrics.

2
New cards
What is the Clever Hans effect, and why is it dangerous?

The Clever Hans effect occurs when a model learns a spurious shortcut instead of the intended semantic concept. Example: a person detector associates yellow safety vests with people and may then miss a person wearing dark clothes or falsely detect an empty yellow jacket.

3
New cards
What is the difference between SC1 and SC2?

SC1: mismatch already present during development — the dataset does not sufficiently represent real-world operating conditions.
→ MA1: systematic data acquisition covering the ODD.

SC2: the distribution changes after deployment over time.
→ MA10: monitoring, collecting new data, retraining and revalidating.

4
New cards
Why are rare critical situations an AI safety concern, and how are they addressed?
Rare but hazardous scenarios form a long tail and are difficult to collect or anticipate completely before deployment. MA5 uses targeted and field testing to find such cases, while MA6 analyzes test results and uncertainty systematically and feeds discovered weaknesses back into development.
5
New cards
What are unreliable confidence and DNN brittleness, and how are they mitigated?

SC5 means model confidence scores can be overconfident or poorly calibrated, so downstream safety functions cannot trust them; MA2 calibrates confidence outputs. SC6 means small changes such as noise, weather, translations or adversarial perturbations can change predictions; MA4 addresses this with realistic threat models and defense methods.

6
New cards
Why is improper train/test splitting dangerous?
If highly correlated data such as consecutive video frames or the same locations occur in both training and test sets, the test result becomes artificially high and does not measure real generalization. MA7 requires spatially and temporally separated datasets and documented partitioning rules.
7
New cards
Why is labeling quality safety-critical?
Incorrect or inconsistent annotations directly affect what the model learns and can also distort evaluation results. MA8 therefore uses detailed task-specific labeling guidelines, safety-relevant metadata, multiple annotators or audits, and quality-control procedures.
8
New cards
Why are ordinary metrics such as average mAP insufficient for safety-critical perception?
Average metrics treat errors similarly even though their safety consequences differ. For example, missing a pedestrian close to the vehicle is more critical than missing one far away. MA9 therefore uses safety-aware or risk-weighted evaluation metrics that account for the consequence and context of errors.