Testing and Evaluation

Module 2: Important Characteristics of Assessments

Validity
“Validity is an evaluation of the adequacy and appropriateness of the interpretations and uses of assessment results” (Linn & Gronlund, 2000, p.73).  Does a particular reading comprehension test tell us which students have good reading comprehension skills and which do not?  Does the test measure what it intends to measure (BC Teachers’ Federation, 2003)?  It is important to make sure that assessments are used for the specific purpose and target population for which they were designed.  If not, their use will not be valid.  For example, if the results from a test designed to measure vocabulary at the elementary school level were used as a measure for reading comprehension, those results would not be valid.  Administering that assessment to middle school students as a measure of vocabulary would not be valid either.

Reliability
Reliability refers to whether scores are consistent and dependable over time.  An assessment is said to be reliable when its results are consistent and the level of measurement error is small.  Assessment results cannot be expected to be totally consistent or reliable.  They actually represent a measure of a sample of performance at a particular time.

Many factors other than what is being assessed may affect assessment results.  For example, if a test is measuring writing skills, what factors other than students’ writing skills may affect the results?  Examples might include variations in effort, attention to task, familiarity with the test items and topics students are being asked to address, and who is scoring the test.  To go a step further, would performance have differed if the student were assessed on a different day, with a different sample of items or if a different rater or teacher scored the test (Linn & Gronlund, 2000; BC Teachers’ Federation, 2003)?

Reliability, or consistency of assessment results, is necessary for validity to be possible.  An assessment that yields inconsistent (unreliable) results cannot produce valid information about what is being measured (e.g., phonics skills). However, high reliability of assessment results means that results are consistent, but does not necessarily mean you are measuring what you intended to measure or that you are using the results appropriately.  For example, an assessment may produce consistent results but may not appropriately measure the subject area or domain it is intended to measure (Linn & Gronlund, 2000).  It is akin to measuring a person’s blood pressure with an inaccurate sphygmomanometer or blood pressure monitor.  You may get consistent results over time, but they are not a valid or accurate measure of that person’s blood pressure.

In This Week

Participants will

Week 1