ΨEmpiralis
← Back to Learn

Reliability and Validity, Explained Simply

Two concepts decide whether a psychological test is any good. What they mean -- explained with a bathroom scale instead of formulas.

Reliability: Does the test measure precisely?

Reliability describes how consistently a measurement tool works, regardless of whether it's measuring the right thing at all. A bathroom scale that shows 150 lbs and then 163 lbs on two weigh-ins seconds apart is unreliable, even if one of those numbers happens to be correct by chance. For psychological questionnaires, reliability is usually reported as Cronbach's alpha -- a value between 0 and 1 that expresses how consistently a test's individual questions hang together. Values above .80 are generally considered good; most tests on Empiralis fall at or above that range, which we reference on each test's own info page.

Validity: Does the test measure the right thing?

Validity is the more important, but harder to verify, property: does the test actually measure the construct it claims to? A bathroom scale can be extremely precise and still be completely useless for answering "how happy am I?" -- it's reliable, but not valid for that purpose.

For psychological tests, validity is checked partly by whether results correlate with other, already-established measures the way the underlying theory predicts. The PHQ-9, for instance, correlates strongly with clinical depression diagnoses from structured clinical interviews -- one reason it's used in primary care worldwide.

Why both matter together

A test can be reliable but not valid (the precise but wrongly calibrated scale). It can also be theoretically well-designed but measure unreliably due to poorly worded items. Only when both are sufficiently present does a questionnaire produce meaningful results -- and that's exactly what the validation studies behind every Empiralis test are checking.

Advertisement

Related tests