Scientific Background
Every test on Empiralis is based on freely available, peer-reviewed questionnaires. For each test we document the original source, its license status, and what studies say about its strengths and limits.
What does "science-based" actually mean?
Psychological questionnaires are developed and evaluated against established psychometric criteria: reliability (measurement precision), validity (does the test actually measure what it claims to?), and norming (comparability against a reference sample). An online self-test can only partially meet these criteria, since e.g. the conditions under which it's taken aren't standardized. Our tests are meant for self-reflection, not diagnosis.
Articles
The methodology behind the tests, explained in plain language -- what separates a validated instrument from a random internet quiz.
What Actually Makes a Test Scientific
Between a carefully developed questionnaire and a random internet quiz often lie years of research and thousands of participants. A look behind the scenes -- in plain language, with real examples.
What “above average” means here — and what’s still missing
A norm is always relative to a reference group. We show which sample your result is compared against, where we use our own orientation ranges, and why we don’t (yet) run our own validation research.
How Reliable Are Online Personality Tests, Really?
Between serious, decades-researched questionnaires and pure entertainment quizzes lie entire worlds. How can you tell the difference?
Reliability and Validity, Explained Simply
Two concepts decide whether a psychological test is any good. What they mean -- explained with a bathroom scale instead of formulas.
When Is a Screening Result a Reason to Seek Professional Help?
A concerning test result isn't a verdict -- but it's not a coincidence to ignore, either. Orientation, without causing alarm.
Personality Test vs. Clinical Screening -- What's the Difference?
Both consist of questions with answer scales -- but they measure fundamentally different things and carry different consequences for you.
Guides
Background, concrete action tips, and a printable checklist and note table for select topics.
Depression: Information, Action Tips & Help
What a depressive episode actually is, what research shows genuinely helps, and where to find support.
Anxiety & Anxiety Disorders: Information, Action Tips & Help
Where normal worry ends and a treatable anxiety disorder begins -- and what research shows genuinely helps.
Well-Being & Life Satisfaction: Information & Action Tips
What research on psychological well-being actually shows -- and what's easy to bring into everyday life.
Alcohol & Substance Use: Information, Action Tips & Help
Between low-risk drinking, hazardous use, and dependence -- where are the lines, and what helps to course-correct early?
Eating Behavior & Eating Disorders: Information, Action Tips & Help
When worry about eating is more than a phase -- and why speaking up early makes the biggest difference.
Burnout: Information, Action Tips & Help
More than exhaustion after a hard day -- and why acting early makes the difference.
Attachment Style & Relationships: Information, Action Tips & Help
How attachment patterns form, why they aren't a fixed fate, and what supports changing them.
Loneliness: Information, Action Tips & Help
Why loneliness is a feeling, not a fact -- and what genuinely helps you feel connected again.
Adult ADHD: Information, Action Tips & Help
Why ADHD in adults is often recognized late -- and what genuinely eases daily life.
Autism Spectrum in Adults: Information, Action Tips & Help
Why autistic traits in adults -- especially when well-masked -- are often recognized only late.
Parental Stress & Family Function: Information, Action Tips & Help
Parenting means joy and strain at once -- why easing the load early helps both parents and kids.
Building Self-Esteem: Information & Action Tips
What self-esteem actually is, and why it's trainable -- despite what you might assume.
Practicing Self-Compassion: Information & Action Tips
What self-compassion actually means -- and why it's the opposite of self-pity or letting yourself off the hook.
Child Development & Behavior: Information, Action Tips & Help
When parents worry about their child: how to make sense of what you're seeing and take the right first step.
Media Use & Screen Time: Information, Action Tips & Help
Why how a child uses media says more about their well-being than raw screen time -- and what research shows actually helps.
Scientific background by test
Depression
The PHQ-9 was developed by Spitzer, Kroenke, and Williams (1999) based on the DSM-IV criteria for major depression and is one of the most extensively studied screening instruments in clinical psychology and primary care. Numerous studies show good internal consistency (Cronbach's alpha usually above .85) as well as good sensitivity and specificity for detecting depressive disorders. The PHQ-9 is a screening and monitoring tool, not a diagnosis by itself — that always requires a clinical conversation. The original study included 6,000 patients across eight primary care and seven obstetrics-gynecology clinics; criterion validity was additionally checked in 580 people against an independent structured interview by a mental health professional. At the cutoff of 10 points, the PHQ-9 reached a sensitivity of 88% and a specificity of 88% for detecting major depression -- a well-balanced trade-off between missed cases and false alarms.
Anxiety
The GAD-7 was developed by Spitzer, Kroenke, Williams, and Löwe (2006) and is — alongside the PHQ-9 — one of the most extensively studied and widely used screening instruments in clinical psychology. Studies show good internal consistency (Cronbach's alpha usually above .89) and good sensitivity and specificity, not only for generalized anxiety disorder but also as a general indicator for panic disorder, social phobia, and PTSD. The GAD-7 is a screening and monitoring tool, not a diagnosis by itself. In the original validation study on a large primary care sample, a cutoff of 10 points gave the best balance of sensitivity (89%) and specificity (82%). Importantly, these figures don't automatically transfer to every population: a later study in a psychiatric sample (people already in treatment for another mental health condition) found notably weaker discrimination at the same cutoff (sensitivity 74%, specificity 54%) -- a reminder that screening accuracy depends on the population being tested, not just the questionnaire itself.
Well-Being
The WHO-5 was introduced in 1998 by the WHO Regional Office for Europe and, with over a thousand validation studies in more than 30 languages, is one of the best-studied short instruments for psychological well-being. It's used, among other things, in primary care as the first step of a two-stage depression screening: scores of 50% or below are considered a sign of reduced well-being and, per common clinical practice, warrant further evaluation for depressive symptoms. Since 2024, the WHO itself holds the copyright to the WHO-5 and explicitly makes it available worldwide free of charge. The systematic literature review by Topp and colleagues (2015) evaluated 213 studies from 30 countries and consistently confirms good internal consistency (Cronbach's alpha usually above .80), as well as remarkably high sensitivity to change for a five-item instrument -- for example, over the course of treatment. That responsiveness to change, not just raw accuracy as a depression screener, is considered in the literature to be one of the WHO-5's biggest strengths compared to similarly short instruments.
Burnout (CBI)
The CBI was developed in 2005 by Kristensen, Borritz, Villadsen, and Christensen as a freely available alternative to the commercial Maslach Burnout Inventory. It shows good internal consistency (Cronbach's alpha between .85 and .91) and correlates strongly with established burnout and exhaustion measures. The authors emphasize that burnout is not a stable personality trait but can change meaningfully over time -- retaking the test periodically can help make those changes visible. The Personal Burnout subscale (items 1-6) is also discussed in the research literature as a possible screening signal for autistic burnout. The authors themselves published normative data in 2004 (Borritz & Kristensen, nfa.dk): the only clinical cutoff they publish is a score of 50 or more as "a high degree of burnout" (in the PUMA project's comparison sample of roughly 1,900 employees in social/care professions, 16-22% reached this level depending on the subscale). The lower bound of the "average" range was additionally derived from the means/standard deviations reported there: Personal Burnout from a representative Danish population sample (N=1,498, M=32.7, SD=15.7), Work- and Client-related Burnout from the PUMA sample (Work: M=33.0, SD=17.7, N=1,910; Client: M=30.9, SD=17.6, N=1,752), since no general-population data exists for those two work-specific subscales.
Self-Compassion
The Self-Compassion Scale was developed in 2003 by Kristin Neff (University of Texas at Austin) and is the most widely used instrument in self-compassion research worldwide. Self-compassion comprises three complementary pairs of opposites: self-kindness versus self-judgment, experiencing common humanity versus isolation, and mindfulness versus over-identification with negative feelings. Studies consistently show good internal consistency across the six subscales, along with robust associations with lower depression and anxiety and higher well-being. The bands used here were derived from the means/standard deviations reported in the original study's Table 4 (Neff, 2003, N=391 U.S. college students): Self-Kindness M=3.05/SD=0.75, Self-Judgment M=3.14/SD=0.79, Common Humanity M=2.99/SD=0.79, Isolation M=3.01/SD=0.92, Mindfulness M=3.39/SD=0.76, Over-Identification M=3.05/SD=0.96 (each on the 5-point item scale). The "average" range corresponds roughly to one standard deviation around that mean for each subscale. Kristin Neff's own materials explicitly note that there are no official clinical norms for the SCS -- these bands are a reference point, not a diagnostic scheme. There's an active scientific debate about how the SCS should be scored. Muris and Petrocchi (2017) argue that combining positive items (self-kindness, common humanity, mindfulness) and reverse-scored negative items (self-judgment, isolation, over-identification) into a single overall number is misleading, since the negative items correlate more strongly with anxiety and depression (r = .47 to .50) than the positive items do (r = -.27 to -.34) -- meaning a combined score may just be tracking self-criticism rather than self-compassion as a distinct, protective quality. Neff and colleagues (2017) counter that a bifactor analysis across four samples shows a single overall self-compassion factor accounts for the great majority of item variance. This site sidesteps the debate rather than taking a side: it reports all six subscales separately below rather than collapsing them into one combined score, so you can judge the positive and negative sides of self-compassion independently.
Life Satisfaction
The Satisfaction With Life Scale was developed in 1985 by Ed Diener and colleagues as part of research into subjective well-being. It's considered the standard instrument for measuring global, cognitively-evaluated life satisfaction, and is deliberately distinguished from the emotional component of well-being (positive/negative affect). The scale shows high internal consistency across numerous studies and stable test-retest reliability appropriate for a trait measure, while still being sensitive to change -- for example, over the course of therapeutic interventions. The bands below follow the interpretation ranges Diener himself publishes ("Understanding SWLS Scores"), based on the scale's 5-35 point range. The original study reported a Cronbach's alpha of .87 and a two-month test-retest correlation of .82. A meta-analysis pooling 60 studies confirms good internal consistency on average (alpha ≈ .78) across a very wide range of countries, languages, and age groups -- evidence that these five, deliberately broad items hold up robustly across cultural contexts, which isn't a given for such a short instrument.
Family Function
The Family APGAR was developed in 1978 by family physician Gabriel Smilkstein to quickly assess family function in a primary care setting -- named after the well-known APGAR score from obstetrics. The name is an acronym for the five areas it covers: Adaptation (support), Partnership (communication), Growth, Affection, and Resolve (shared time/connectedness). Studies consistently show high internal consistency (Cronbach's alpha .80-.90) and good test-retest reliability. Factor analyses suggest the five items mostly capture a single, global dimension of "satisfaction with family" rather than five separate constructs. The validation study by Smilkstein and Ashworth (1982) compared families with and without known functional problems and found clear differences in the expected direction -- since then the Family APGAR has been used in primary care worldwide and translated into numerous languages. We're not aware of a large-scale, independently published English-language norming study with its own population norms beyond the original development samples; the scoring here follows the internationally established structure and interpretation of the original instrument.
PSC-17 (Parent)
The PSC-17 is a short form developed by Gardner and colleagues (1999) from the original, 35-item Pediatric Symptom Checklist by Jellinek and Murphy (1988). In a comparison study against the much longer Child Behavior Checklist (CBCL) with 269 children, the PSC-17 showed comparable detection accuracy at a fraction of the length (Gardner et al., 2007). The three subscales correlate with different clinical presentations: attention with ADHD, internalizing symptoms with anxiety and depressive disorders, externalizing symptoms with oppositional and aggressive behavior. Like any brief screener, the PSC-17 does not replace a thorough evaluation -- its sensitivity is around 42-73% depending on the condition. How strong these numbers look depends heavily on the comparison standard used: the PSC-17 total score correlates with the CBCL total score at r = -0.60 in Gardner et al.'s (2007) comparison study. Against a broader criterion -- any psychosocial impairment, rather than one specific diagnosis such as ADHD or depression -- other reviews report considerably higher figures (mean sensitivity .90, mean specificity .85) than the 42-73% cited above for individual conditions. In short: the PSC-17 is good at flagging that something may be going on, but less precise at predicting which specific diagnosis it is.
Parental Stress
The Parental Stress Scale was developed in 1995 by Berry and Jones as a shorter alternative to the 101-item Parenting Stress Index. Its distinctive feature: it deliberately captures not just demanding aspects, but also positive, enriching aspects of parenting -- both contribute to the overall picture of parental stress. Higher scores are associated in studies with lower parental sensitivity, more child behavior problems, and lower quality of the parent-child relationship -- the scale is used, among other things, to track change from parenting programs or family support. The bands used here follow a population-based Norwegian sample of 1,096 parents of one-year-olds (Nærde & Hukkelberg, 2020, PLOS ONE), which reports a mean of 31.0 (SD=7.27) on the 18-90 raw score range: the "average" range corresponds to roughly one standard deviation around that mean. In the original study, the total score's internal consistency was Cronbach's alpha = .83, with six-week test-retest reliability of .81. At the level of the two content areas -- demanding versus rewarding aspects of parenting -- internal consistency is considerably less consistent across follow-up studies (roughly .61 to .79 depending on the sample and area) than for the total score. That's a good reason to treat the total score as the more dependable figure, rather than splitting results into the two sub-areas.
Vanderbilt (Parent, ADHD)
The Vanderbilt scales were developed by Mark Wolraich and colleagues based on DSM-IV criteria and validated in a large sample of referred children (Wolraich et al., 2003). They're among the most widely used ADHD screening instruments in U.S. pediatric primary care. The validation study reports very high internal consistency (Cronbach's alpha between .91 and .94 per subscale) and test-retest reliability above .80. When checked against a combination of teacher ratings and a diagnostic interview, the parent-report scale alone reached a sensitivity of .80 and a specificity of .75 -- but a positive predictive value of only .19. In practice, that means most children flagged as elevated by the parent form alone don't end up confirmed on full evaluation. That's precisely why the original scale is never recommended for diagnosis on its own -- it needs to be combined with other sources such as a teacher rating and a clinical interview (see the note below).
Alcohol Use
The AUDIT was developed by the World Health Organization as part of an international multi-center study (Saunders et al., 1993) and is today the most widely used screening instrument for hazardous drinking in primary care worldwide. It covers three areas: consumption (questions 1-3), dependence symptoms (questions 4-6), and alcohol-related problems (questions 7-10). Numerous international validation studies confirm good sensitivity and specificity for detecting hazardous use. The development sample included 1,888 primary care patients (36% non-drinkers, 48% drinking patients, 16% with alcohol dependence). A review across six AUDIT studies reports, at the cutoff of 8, a sensitivity of 97% and specificity of 78% for detecting hazardous drinking, and a sensitivity of 95% and specificity of 85% for detecting harmful drinking -- the test correctly flags the great majority of true cases while producing relatively few false alarms.
GAS-7 (Gaming)
The Game Addiction Scale was developed in 2009 by Lemmens, Valkenburg, and Peter for adolescents, based on classic addiction criteria (following Griffiths, among others): salience, tolerance, mood modification, relapse, withdrawal, conflict, and problems. The 7-item short form used here was validated in 2016 by Khazaal and colleagues in large French- and German-speaking samples from Switzerland and showed good internal consistency. Problematic gaming has been listed as "Gaming Disorder" in the WHO's ICD-11 since 2019 as its own diagnosis -- though that requires marked impairment over at least 12 months and a professional evaluation, which this brief screen cannot replace. Khazaal and colleagues' 2016 validation study tested the short form on a very large Swiss sample (3,318 French-speaking and 2,665 German-speaking men) and found satisfactory internal consistency (Cronbach's alpha = .85) and a single-factor structure in both language groups. One thing worth noting: that validation sample was exclusively male -- how well these figures generalize to women or non-binary people wasn't examined in that study.
Eating Habits
The SCOFF was developed in 1999 by Morgan, Reid, and Lacey in a British primary-care sample and published in the BMJ. In the original study, a threshold of two or more "yes" answers correctly identified roughly 85-100% of people with a diagnosed eating disorder (sensitivity), with similarly high specificity. SCOFF has since been translated into numerous languages and validated internationally. As a pure brief screen, it can produce false positives and false negatives and doesn't replace a thorough evaluation. A systematic review and meta-analysis pooling 25 studies (2020) confirms this good detection accuracy holds up across many different samples: pooled sensitivity of 86% and specificity of 83%, with an overall discrimination (AUC) of 0.91 -- a value generally considered very good in diagnostic accuracy research. These pooled figures reflect a wider range of real-world settings than the original single-sample British study the SCOFF was first validated on.
Attachment Style
The ECR-R was developed in 2000 by Fraley, Waller, and Brennan using item response theory, based on the original ECR by Brennan, Clark, and Shaver (1998), and is considered one of the psychometrically most robust instruments in attachment research. The combination of low attachment anxiety and low attachment avoidance is referred to in the research literature as a "secure attachment style"; high scores on one or both dimensions correspond to anxious-insecure, avoidant-insecure, or anxious-avoidant attachment patterns. The original validation study reports good internal consistency for both dimensions (Cronbach's alpha around .86 for attachment anxiety, .88 for attachment avoidance) and a striking degree of temporal stability for a self-report personality measure: both dimensions remained about 85-86% stable over three- to six-week intervals. That suggests the ECR-R captures a genuinely enduring attachment pattern rather than a passing mood.
Relationship Satisfaction
The Couples Satisfaction Index (CSI) was developed in 2007 by Funk and Rogge (University of Rochester) using item response theory. The goal was more precise measurement of relationship satisfaction than older scales such as the Dyadic Adjustment Scale or the Marital Adjustment Test, both of which the CSI outperformed in direct comparisons. There are 4-, 16-, and 32-item versions; despite its brevity, the CSI-4 short form used here achieves measurement quality comparable to the longer versions and showed high internal consistency (Cronbach's alpha around .94) in the original study. The questionnaire measures a single, global construct — the subjective overall evaluation of one's relationship — not individual facets like communication or sexuality. Funk and Rogge's development study drew on a very large online sample of 5,315 people and compared eight established satisfaction scales using item response theory. It found that established, much longer scales like the Dyadic Adjustment Scale measure imprecisely at many points along the satisfaction spectrum despite their length -- the CSI was deliberately built to discriminate more precisely across the full range of scores. The full 32-item version reached an exceptionally high internal consistency (Cronbach's alpha = .98); the CSI-4 short form used here deliberately trades some of that precision for brevity, without losing much measurement quality in the process.
Social Support
The MSPSS was published in 1988 by Gregory Zimet, Nancy Dahlem, Sara Zimet, and Gordon Farley in the Journal of Personality Assessment. It consists of twelve statements that group into three four-item subscales: support from a special, close person, from family, and from friends. In the original study the scale showed good internal consistency (Cronbach's alpha .88 for the total scale; .91 / .87 / .85 for the three subscales) and a stable three-factor structure that has since been replicated across many languages and samples. Scoring is by means: for each subscale, the sum of its four items divided by four; for the total score, the sum of all twelve items divided by twelve. The original 1988 study (N=275 university undergraduates) also reported test-retest reliability of .72 (significant other), .85 (family), and .75 (friends) over up to three months, and found the expected pattern of convergent validity: higher perceived social support was associated with lower depression and anxiety symptoms (measured via the Hopkins Symptom Checklist) and with higher well-being, positive affect, and coping -- exactly the kind of external validation you'd want from a support scale that claims to matter for mental health.
Loneliness
The UCLA Loneliness Scale was originally developed in 1978 by Russell, Peplau, and Ferguson, and revised by Daniel Russell into the third version used here in 1996. It's the most widely used research instrument in the world for measuring subjective loneliness and shows very good internal consistency across numerous studies (Cronbach's alpha usually above .90). Importantly: the scale measures the subjective feeling of loneliness, not the objective number of social contacts -- you can feel lonely in a crowd, and be alone without feeling lonely. The bands used here are based on the reference values reported in the original study (Russell, 1996, Table 2) for two adult samples using the full 20-item scale (college students, N=487, M=40.08, SD=9.50; nursing staff, N=305, M=40.14, SD=9.52): the "average" range corresponds to roughly one standard deviation around that pooled mean (M≈40.1, SD≈9.5). The original study does not define a formal clinical cutoff -- the scale is meant to be interpreted dimensionally, not categorically.
Big Five
This test is based on the International Personality Item Pool (IPIP), a freely available collection of personality items developed as a public alternative to commercial questionnaires like the NEO-PI-R. The 50 items used here correspond to the well-known IPIP representation of the five factors after Costa & McCrae (Goldberg, 1992) -- specifically the IPIP's 10-item scale for each factor. This is for self-reflection, not diagnostic classification. Studies on the Five-Factor Model show robust international replication of the factor structure, as well as good test-retest reliability over periods of several years. International validation studies report satisfactory internal consistency for all five scales for a freely available research instrument (Cronbach's alpha typically between .73 and .84), along with good agreement with established commercial questionnaires like the NEO-FFI. That agreement isn't equally strong across all five dimensions, though: it's highest for Conscientiousness, Extraversion, and Emotional Stability, and noticeably lower for Agreeableness and Openness to Experience -- a known weak spot that shows up in other Big Five questionnaires too, and worth keeping in mind when interpreting those two scores specifically.
Dark Tetrad (SD4)
Paulhus and Williams coined the term "Dark Triad" in 2002 for three socially unpopular but overlapping personality traits -- narcissism, Machiavellianism, and psychopathy -- which proved to be distinct but related constructs. Buckels, Jones, and Paulhus (2013) showed that everyday sadism (enjoyment of cruelty toward others) represents a fourth, empirically separable facet -- since then, researchers have spoken of the "Dark Tetrad." In 2020, Paulhus, Buckels, Trapnell, and Jones published the SD4 (Short Dark Tetrad) in the European Journal of Psychological Assessment as a compact, 28-item short form that captures all four traits with good internal consistency (Cronbach's alpha between .76 and .81), replacing the previous three-trait SD3 short scale from the same author team. Importantly: these traits are understood as normally distributed personality characteristics in the general population ("subclinical") -- they are conceptually distinct from the clinical diagnoses of Narcissistic or Antisocial Personality Disorder, or clinical psychopathy (which is assessed via, e.g., the clinician-administered PCL-R). The bands used here are based on the item means and standard deviations reported in the original study for a sample of 660 University of Winnipeg students, and serve only as a rough point of reference, not a clinical norm.
Self-Esteem
The scale was developed in 1965 by sociologist Morris Rosenberg and remains the most widely used instrument for measuring global self-esteem in psychological research today. Numerous studies across many countries and languages show very good internal consistency (Cronbach's alpha usually above .85) and stable test-retest reliability. The scale deliberately measures a one-dimensional, global sense of self-worth, not individual facets like a sense of competence or body image. Methodological research has long debated whether the scale truly measures just one dimension: because five of the ten items are positively worded and five negatively, some studies find two separate factors -- "self-confidence" and "self-deprecation" -- rather than a single factor, likely because negatively worded items get answered somewhat differently regardless of a person's actual self-esteem. More recent studies using bifactor models confirm that underneath this wording effect, a single, substantial global factor remains -- which is what the scoring here reflects.
ASRS (Adult ADHD)
The ASRS was developed in 2005 by a WHO working group together with Ronald C. Kessler (Harvard Medical School) and colleagues, and is one of the most widely used screening instruments internationally for adult ADHD. The full version has 18 items; the six items used here ("Part A") were identified in the original study as the most statistically informative of all 18, and are validated as a standalone brief screener. Importantly: this questionnaire checks for current symptoms, but doesn't replace the evidence, required for an ADHD diagnosis, that similar difficulties were already present in childhood. In the development sample, the six-item short screener reached a specificity of 99.5% at moderate sensitivity of 68.7% using a threshold of 4 or more elevated answers -- meaning the test misses a meaningful share of true cases, but very rarely produces a false positive. Adler and colleagues (2006) report good internal consistency (Cronbach's alpha of approximately .88). Validation studies of the successor instrument, the ASRS-5, find comparable reliability but a different sensitivity-specificity balance depending on language and sample -- a reminder that these numbers can shift somewhat depending on the population studied.
CAT-Q (Masking)
The CAT-Q was developed in 2019 by Laura Hull and colleagues at University College London, after clinical observations showed that many autistic people -- particularly women and adults diagnosed later in life -- camouflage their traits so successfully that classic screening instruments like the AQ don't reliably capture them. Camouflaging is associated with increased mental exhaustion, anxiety and depression, and delayed diagnosis. The questionnaire shows good internal consistency for all three subscales in validation studies and reliably distinguishes between autistic and non-autistic samples. Important context: a 2026 critical review of 389 studies (Arnold, Bitsika & Sharpley, 2026, Autism) credits the CAT-Q with excellent reliability but finds mixed evidence for its validity -- total scores correlated more strongly with social anxiety than with autistic traits already in Hull's own 2019 validation study, and more strongly with ADHD traits than autistic traits in a separate study (Ai et al., 2023). This limits the instrument's specificity to autism, which is why the review explicitly cautions against using it as a screening tool. Several non-English translations (including Dutch, Japanese, Swedish, and Traditional Chinese) also failed to replicate the original three-factor structure, and the Dutch and French versions failed to show measurement invariance between autistic and non-autistic groups -- a sign that results aren't automatically comparable across languages and cultures. The CAT-Q's own authors now explicitly caution against diagnostic use as well: a 2026 Perspective article (Hannon, Hull, Lai, Magiati & Mandy, 2026, Autism in Adulthood) -- with CAT-Q co-creator Laura Hull as a co-author -- recommends against using CAT-Q scores in diagnostic decision-making or as a treatment-outcome measure.
Autistic Traits
The RAADS-R was published in 2011 by Ritvo and colleagues as a revised version of the original RAADS, and is one of the most widely used screening instruments in research and clinical practice for autistic traits in adults. It's organized into four domains: social relatedness, circumscribed interests, language, and sensory-motor. The original study proposed a total-score threshold of 65 points, above which autistic traits are likely (non-autistic control group mean: approximately 21 points). Importantly: both lower and higher scores occur in both autistic and non-autistic people -- in particular, pronounced "masking" (see the CAT-Q) can lower the score. This test can therefore never replace a diagnosis, only provide an indication that further professional evaluation may be worthwhile.
Resilience
The Brief Resilience Scale was developed by Bruce W. Smith and colleagues in 2008, published in the International Journal of Behavioral Medicine. The authors deliberately designed it to fill a gap they identified in existing resilience research: many older resilience scales actually measure resources associated with resilience (e.g. social support, optimism, sense of purpose) rather than the ability to recover itself, making it hard to tell whether someone is resilient or simply has access to more resources. The BRS was validated across four samples (undergraduate students, cardiac rehabilitation patients, people with fibromyalgia, and a community sample of women) and has since become one of the most widely used brief resilience measures in health psychology research, showing good internal consistency and correlating with better physical and psychological health outcomes after adversity. Across those four original samples, internal consistency ranged from Cronbach's alpha .80 to .91, and all items loaded onto a single factor. Test-retest reliability was more moderate and declined somewhat over time (ICC = .69 over one month, .62 over three months) -- a pattern consistent with resilience being at least partly a state that can shift with circumstances, not a completely fixed trait. That's worth keeping in mind if you retake this test later and get a different score: some genuine change over time is expected, not necessarily a measurement error.
SOMEDIS-A (Social Media)
SOMEDIS-A was developed in 2021 by Paschke, Austermann, and Thomasius at the German Center for Addiction Research in Childhood and Adolescence (DZSKJ, University Medical Center Hamburg-Eppendorf) -- adapted from the same team's already-validated Gaming Disorder Scale for Adolescents (GADIS-A). It is the first screening instrument for problematic social media use built directly on the ICD-11 criteria for Gaming Disorder, rather than the DSM-5 Internet Gaming Disorder criteria used by older scales. It was validated in a representative German sample of 931 parent-adolescent pairs (adolescents aged 10-17), recruited through the market research institute forsa. The two-factor structure (cognitive-behavioral symptoms, negative consequences) showed good to excellent internal consistency and good criterion validity (associations with the Social Media Disorder Scale, the PHQ-9, and the Perceived Stress Scale). Specifically, the validation study reports excellent internal consistency for the total score (Cronbach's alpha = .91) and moderate but statistically significant associations with depressive symptoms (PHQ-9: r = .48) and perceived stress (PSS-10: r = .43) -- a sign that problematic social media use goes along with, but isn't the same thing as, depression or stress. In the representative sample, 6.3% of adolescents met the criteria for problematic use under this instrument.
Functional Impairment
The WFIRS-S was developed by Margaret D. Weiss to complement pure symptom checklists like the ASRS-v1.1 with a question they don't answer: how much do these difficulties actually affect day-to-day functioning? A conceptual review (Weiss et al., 2018) summarizes that symptom improvement doesn't automatically track with functional improvement -- the two are related but distinct. The WFIRS-S has been translated into 18 languages and psychometrically evaluated in student and adult samples in the US, Japan, and Iran, among others, with consistently good internal consistency. Unlike the parent-report version (WFIRS-P), no study to date has established a validated cutoff for the self-report form -- Weiss et al. (2018) explicitly name this as open research. What is available are reference means from individual samples (e.g. Canu et al., 2016: adult students screening positive for ADHD, M=0.88; screening negative, M=0.35). Official scoring uses the mean per domain, not the sum, so items marked "not applicable" don't distort the calculation. Canu and colleagues' 2016 study, in a very large sample of 2,093 college students, reports exceptionally high internal consistency (Cronbach's alpha above .96 for both self-report and informant-report) and good agreement between the two report forms. The same study also found an important caveat: measured impairment correlated not only with ADHD symptoms but also with internalizing symptoms like anxiety and low mood. The WFIRS-S measures general everyday impairment, not ADHD-specific impairment -- an elevated score can have causes other than ADHD.
Functional Impairment (Child)
The WFIRS-P was developed by Margaret D. Weiss, Michael B. Wasdell, and Melissa M. Bomben and published in 2005 as part of the Canadian ADHD Practice Guidelines. It complements pure symptom checklists (like the Vanderbilt parent scale) with a central, often-overlooked question: how much do these difficulties actually affect day-to-day functioning? A conceptual review (Weiss et al., 2018) summarizes that symptom improvement doesn't automatically track with functional improvement -- a child can be symptomatically better and still notably impaired in specific life areas, or the reverse. The WFIRS-P has since been translated into 18 languages and validated in various clinical samples, with consistently good to very good internal consistency. Official scoring uses the mean per domain, not the sum, so items marked "not applicable" don't distort the calculation. A large pooled validation (Gajria et al., 2015) combined data from seven randomized controlled trials across 2,357 children and adolescents with ADHD (ages 6-17, mostly from Europe and North America) and found Cronbach's alpha above .70 for every domain, along with confirmed responsiveness to treatment. One domain is a partial exception worth noting: "Risky Activities" showed somewhat lower test-retest reliability across studies -- a plausible reason is that risky behavior (like dangerous behavior in traffic) occurs less often and less regularly in children than, say, school or family difficulties, which makes the scale statistically less stable there. A bit more caution is warranted when interpreting that one domain specifically.
Substance Use (Teens)
The CRAFFT was developed by the Center for Adolescent Substance Use Research (CeASAR) at Boston Children's Hospital and is one of the best-validated brief screens for risky substance use in adolescence. In a re-evaluation against DSM-5 criteria (Mitchell et al., 2014), the proportion of people with a diagnosable substance use disorder rose sharply with the number of "yes" answers -- from 32% at one "yes" answer to over 90% at four or more. A score of 2 or more is considered a sign of a problem serious enough to warrant further evaluation. In the original validation study by Knight and colleagues (2002), a threshold of 2 or more proved to be the optimal balance for detecting a diagnosable substance use disorder in 14- to 18-year-olds (sensitivity 80%, specificity 86%). For the more specific question of dependence on alcohol or other substances, the same threshold performed even better in follow-up studies (sensitivity 92%, specificity 80%) -- the CRAFFT is especially reliable at catching the more severe cases.
KINDL (Children 7-13)
The KINDL was developed in 1994 by Prof. Monika Bullinger and revised into the KINDL-R in 1998 by Prof. Ulrike Ravens-Sieberer and Bullinger. It's one of the most established instruments internationally for measuring health-related quality of life in children and adolescents, translated into more than 30 languages and used in numerous national and international studies. The six-factor structure (physical, emotional, self-esteem, family, friends, school) has been confirmed across many validation studies. Important: the official KINDL manual itself doesn't define severity categories like "low/typical/high" -- its own scoring compares the 0-100 scale scores against age- and sex-specific population norms, which exist for Germany (from the nationally representative KiGGS study) but not for a comparable US or English-speaking sample. The three-tier breakdown used below is Empiralis's own simplified, age-independent orientation aid, not an official category of the instrument, and not a population comparison. A psychometric evaluation of the English-language version among English-speaking Singaporean children (Wee et al., 2007) found internal consistency ranging from acceptable to good depending on the subscale and age group (alpha .40-.71 for the 7-13 Kid version, .44-.84 for the 14-17 Kiddo version) -- a reminder that, especially at younger ages, individual subscales with only four items each can be less statistically stable than the total score. A separate large-scale psychometric analysis of the original German version (Bullinger et al., 2008, the BELLA study) reports good internal consistency for the total score (Cronbach's alpha = .85 in child/adolescent self-report, .86 in parent report) -- parent and self-ratings agree reasonably well overall, even though they aren't identical.