PSY 133 Exam 2 Study Guide

Study Guide, Exam II

Psy 133

 

 

Where does one obtain information about tests: know major sources (MMY and TIP) and type of info that can be found in each.

 

Validity: Chapter 5

         Be able to define validity theoretically and mathematically

Theoretically: the extent to which inferences drawn from an instrument are correct

Mathematically: observed score is X= T+e T= portion of score that is consistent

  • X= V (valid part) + I (irrelevant part) + e

  • T= V+E

What is E? - Error

        

         Relationship to reliability

                     Correction for Attenuation – pg 127-128

for what is it used, know formula for perfect reliability

> used to estimate what the correlation would have been if the variables had been perfectly reliable

>What is the formula for perfect reliability?

                     What is range restriction, and how does it affect correlation coefficients

Range restriction – when there is restricted range of variation in scores, resulting in small changes in scores resulting in much larger differences in correlations or rankings

Types of Validity Evidence (advantages and limitations of each, examples of methods for each):

                     Content Validity: evidence that the content of a test representing the domains its supposed to be representing

                                 Uses to find items representative of domain or cover bases. Often used in achievement tests or licensing

                                 Methods of determining content validity: Lawshe: Content Validity Ratio

  • Compute CVR for each item

    • CVR= (ness – nunness)/ntotalraters

  • Compute CVR for test (content validity index)

    • CVI= SCVR/nitems

 

                     Face Validity (not really true validity): when/why is it desirable/when can be a drawback

How the test appears to test takers as relating to its purpose. Good for wanting to keep them on track, make them feel confident, bad if exam’s purpose is for someone to not know whats being tested like a depression or bipolar screening or something.

                     Construct Validity: observed correlations between the test and other measures provide evidence for the meaning of the test

 

                                 Process of development: hypothesis testing, developing nomological network

                                 Measures

  • Process: accumulate evidence/data about what construct test measures

  • Nomological network: Refers to the network of concepts that the test relates to and doesn’t relate to

 

                                             internal, experimental, correlational data

                                             (criterion related, discriminant, convergent,

                                             factor analysis with varimax rotation) 

Correlation data is referring to convergent and discriminant validity. 

Convergent evidence - looking if measures that measure same construct are more highly inter correlated 

Discriminant evidence  - measures that demonstrate a test measures something different from what other available measures are testing (aka what it does that others can’t)

MTMM: know how to interpret one (convergent validity, discriminant validity, method bias).  In reference to those weird matrix tables 

                     Criterion-Related Validity:how well a test corresponds with a particular criterion 

                                 Concurrent vs predictive (advantages/limitations of each)

Predictive validity evidence - the evidence for the criterion validity where the test forecasts scores on a critieron in the future (ex: ACT and future GPA)

Concurrent validity evidence- evidence for criterion validity where the test and criterion are administer at the same point in time 

                                 Criteria:conceptual vs actual criterion, the gap between them is RELEVANCE 

                                             actual vs conceptual

                                             characteristics of good criteria: deficiency, contamination, relevance

Contamination- stuff picked up on criterion that is irrelevant 

Deficiency - the performance area not tapped into by criterion (what the criterion neglects to include) 

Relevance - the overlap between conceptual and actual criterion 


                                             why important to have good criteria why IS it important?? - Having good criteria minimizes contamination and deficiency. You also might get poorer validity or reliability if your criteria are bad.

                                 Indices of criterion related validity: better decision making than w/o test 

Better decisions than other devices 

Must be stable and generalizable 

Must provide utility 

Level (single reliability) ——-> .4 is strong 

                                   valid coeff, group differences, decision-making accuracy.

Group differences = known group validity. Predicted groups, does the test distinguish between who has a disorder or who does not 

                                             Validity-coefficients

                                                         How is a regression equation used

To make a prediction accurately 

                                                         What is cross-validation     We use it to not lose predictive value (aka prevent shrinkage)

Standard Error of Estimate (know formula and how to create confidence intervals around predicted scores for 68% and 95% confidence intervals – will not be using table)



Multiply 2x  ^^ then + and - to get 95% 

                                             What factors can affect level of validity

Unreliability in measure 

Sample characteristics 

Range restriction 

Base rate

Selection ratio 

                                             What is validity generalization

the extent to which we can generalize validity coefficients across situations. 

                     Interpreting validity coefficients:

                                 What do you need to consider (uses)

Utility = Benefit – Cost  

                                 What is an appropriate validity level.

Level is going to matter by context but usually rxx=.4 raw is acceptable 

                     Comparison of types of validity, be able to give examples of each

 Honestly just study what previously is on here and I think you should have this covered. 

Decision Making Accuracy:  what is Utility

Utility = the cost- benefit       AKA → does the cost outweigh the benefits? If yes, no utility. 

         - What are true positives, true negatives, false positives, false negatives?

True positives - have trait, trait detected

False positives - do not have trait, trait detected

True negatives - do not have trait, no trait detected

False Negatives - have trait, not detected 

         -Total Hits (accuracy rate), Positive Hits (detection rate), Base Rate, Selection Ratio

Total Hits = (True positives + true negatives )? (true positives +false positives + true negatives + false negatives) 

Positive Hits (True positives) / (true positives + false positives) 

Base Rate (true positives + false negatives) / (true positives + false positives + true negatives + false negatives) 

Selection ratio (true psoitives + false positives)/(true positives + false positives + true negatives + false negatives) 

                     (know how to calculate these from a table such as in your worksheets)

         - What are the relationships between validity, base rate, selection ratio 

o   Higher validity –hit more accurate decisions

Misses are very small the more valid your data is  

Selection ratio should be higher than base rate if valid 



         -How do these play a part in determining how useful a test is beyond the use of a validity                        coefficient? What are Taylor-Russell tables

Taylor Russel tables - those tables that show the ratio of hits/accurate decisions based on level of validity

If a test is useful, it will result in better decisions 

         - very broadly, what is UtilityAnalysis.

         - how does incremental validity play a part in decisions about test use?

  How much added validity is explained by adding a new predictor. 

If not getting more info with a new predictor, its not necessary. 


Test bias and issues in test administration (Chap 19 & 7):

         Test Bias: form of systematic error, Consists of subgroup differences in test score implications that are not the result of real different in a measured characteristics

Main types: content, criterion-related, internal structure, testing environment/conditions

Content bias - issues based on the content of a test. 2 people w/ same abilities will have different scores 

Internal structure - issues with internal structure of an exam

Testing environment - different environments creating bias during examination 

         Prediction Bias: Differential prediction and differential validity

                     Explain what they are, prevalence, and possible explanations

Differential prediction - Think of prediction/regression lines predicting different outcomes for different groups (example in class: gender and ACT score). think BEST FITTING LINE. 

Differential Validity - the extent to which a test has different meanings for different groups of people, think 2 regression line scores, both with rxx above 0, but significantly different from EACHOTHER. think MAGNITUDE.


         Stereotype Threat:

                     What is it

Stereotyped group preforms to stereotype under certain conditions 


                     How does it impact test scores

Harms performance on exams

Especially if it is important or the difference between groups is made noticeable. 

                     What are methods of reducing

Minimize test importance 

Minimize group differences 

Indicate no difference in performance across groups

         Methods of trying to reduce bias and their effectiveness

-   Change testing

Culture Equivalent Tests (Culture free vs culture fair tests)

                 Culture specific tests: what and why created?  e.g. Chitling,

Intelligence tests created to cater to specific cultures. Created because different cultures emphasized different things and performed differently on exams. 

                 Culture-reduced testing:

Creation of tests that are mostly absent of cultural influence (mostly, because none really achieved 100%)

                             Specific tests:

- Raven's Progressive Matrices

- Cattell's Culture Fair Test of Intelligence


- What is the SOMPA system

- Change testing conditions

- Consider tests within context

- Change environment (SES)

 

Writing test items: Chapter 6, and pages 106-107(IRT) and 127-128 (discriminability analysis)

         Different types of items – issues

         When should you guess on an item, and when should you not guess?  What is "correction for                   guessing"?  You should know formula.

         Item Analysis: Know how to calculate (if relevant) and interpret each of these

                     Item Difficulty (what is optimal level of variability and how to figure it)

                     Item Discrimination

                     Item Distractors

         Item Response Theory:

                     What is it and for what is it used

                     What is an item characteristic curve (be able to describe what an ICC tells about an item)