Mercor SPL

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/43

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:41 AM on 8/19/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

44 Terms

1
New cards

Generation

From scratch, no starting text

2
New cards

Open QA

No single fixed correct answer

3
New cards

Closed QA

One specific, checkable correct answer

4
New cards

Classification

Predefined category/label

5
New cards

Brainstorming

Multiple ideas, not one final answer

6
New cards

Chat

Multi-turn conversation; earlier turns affect later ones

7
New cards

Rewrite

Form/tone/style changed, meaning and length kept

8
New cards

Summarization

Text condensed, key info kept, shorter

9
New cards

Verifiable answer → data type

SFT, ideally with an automated checker (math, code, closed QA)

10
New cards

Judgment call → data type

RLHF / rubric-graded preference data (tone, reasoning quality, open-ended writing)

11
New cards

SFT: how it works

Ideal answer shown, model copies it. Only type with no grader

12
New cards

RLHF / Preference Labeling

Two outputs compared, human picks better, trains a reward model

13
New cards

Rubric-Based Evaluation

Criteria written once by experts and then judge model at scale. The dominant method

14
New cards

RL Environments

Agent acts in a simulated/real setting, rewarded on outcomes, not just text

15
New cards

Wrong format/structure/steps

SFT (demonstrate the right answer)

16
New cards

Capable, but wrong tone/style/character

RLHF (learn human preferences)

17
New cards

Inconsistent quality on complex, subjective reasoning

Rubrics (pinpoint where reasoning fails)

18
New cards

Can't interact with tools/systems beyond a chatbox

RL Environments (measure outcomes of actions)

19
New cards

pass@k

Capability metric: probability of at least one success in k tries

20
New cards

pass^k

Reliability metric: probability all k tries succeed

21
New cards

Hill-climbing sweet spot

Ideal prompt difficulty: ~20–40% model score, not too easy or too hard

22
New cards

Model-as-judge

Rubric written once by humans; model applies it to score at scale

23
New cards

Single-turn eval

One prompt, one response, one grade

24
New cards

Multi-turn eval

Back-and-forth conversation; early mistakes compound

25
New cards

Agent eval

Tool use across many turns, changing state; highest cascading-failure risk

26
New cards

Why evals matter (3 reasons)

Product, scoreboard, directs experts

27
New cards

APEX

Tests high-value professional knowledge work

28
New cards

GDPval

Real tasks across 44 occupations, top 9 US GDP sectors

29
New cards

SWE-bench Verified

Real GitHub issues resolved, patch passes tests

30
New cards

Humanity's Last Exam

Extremely hard, probes near/superhuman capability

31
New cards

HealthBench

Medical/clinical reasoning benchmark, safety-grounded

32
New cards

τ-Bench

Multi-turn simulated agents, retail/airline settings

33
New cards

Benchmark saturation

Top scores near 100%, no longer useful. SWE-bench Verified: ~30% to >80% in months

34
New cards

Open Box: 3 premises

Data is core IP; annotator quality is the biggest lever; speed is critical

35
New cards

Open Box: 5 steps

Propose, source/vet, design pipeline, measure, swap talent

36
New cards

Prompt: the 6 requirements (CFUM SORT)

Critical failure, multi-step, unambiguous, objective, relevant, timeless

37
New cards

Rubric: the 8 requirements (TF SC P1 UTEREA)

T/F, unambiguous, self-contained, one check, expert agreement, timeless, matches prompt, error ranges

38
New cards

Pipeline hierarchy

Campaign > batch > task

39
New cards

IAA (Inter-Annotator Agreement)

Consistency of reviewers grading the same work; Cohen's Kappa = 2 raters, Fleiss' Kappa = more than 2

40
New cards

Stalled batch / failed gate diagnosis

Instructions gap, task difficulty, low performers, pause wave

41
New cards

QC steps & fixes

Writer self-check, calibration session, gold-set spot checks, consolidated feedback, targeted rubric fix, gold-example reference

42
New cards

Causes of bad data

Unrepresentative data, inconsistent labels/low IAA, reward hacking, instructions gap, task difficulty, low performers

43
New cards

Foody's blog thesis

Models are smart but not trained to do the job

44
New cards

Macrosoft: 6 required elements

Data type, domain+why, expert archetype/sourcing, quality mgmt, ramp plan, billing