Psychometric Assessment Records: Documenting Without Distorting

WISC-V, WAIS-IV, MoCA, BDI-II: a score on its own tells you nothing. Version, date, confidence interval, and context are what make a record meaningful.

A score without context is worthless

"IQ 98." Written alone in a file, that single line raises more questions than it answers. Which test? Which version? On what date, and at what age? With what confidence interval? What was the patient's condition that day? Without these details, the number has nothing to be compared against — and worse, it travels: it gets copied into a referral letter, cited by a school, or brought back to challenge the clinician three years later.

Documenting a psychometric assessment therefore means documenting a result along with the conditions under which it was obtained. It is as much a writing task as it is a testing one.

What must accompany every test

  • The instrument and its version: WISC-V, WAIS-IV, WPPSI-IV, K-ABC-II, MoCA, MMSE, BREF, Rey Figure, TMT, Vineland-3, Conners-3, CBCL, BDI-II, STAI;
  • the date of administration and the examiner;
  • the index scores, one by one, with their values;
  • the confidence interval, where provided by the publisher;
  • the normative sample used, along with any relevant caveats — fatigue, language barriers, testing conditions.

A word on caveats: they do not weaken an assessment — they make it honest. A WISC-V administered in French to a child whose schooling has been in Arabic deserves to be read with that in mind.

Test items belong to the publishers

A patient record stores results — scores, indices, percentiles — not test materials. The items, stimulus plates, and scoring tables are protected by their publishers' intellectual property rights, and sharing them would also compromise the validity of the assessments: a test whose items are in circulation no longer measures much of anything.

This is also why a practice management platform has no business "administering" tests. It captures what the psychologist has obtained using their own licensed materials.

Comparing two administrations

The question that always comes up at a follow-up assessment is the same: is the difference meaningful? A four-point gain in Full Scale IQ could reflect a genuine improvement, a learning effect, or simply the test's margin of error.

This is where two often-overlooked elements become essential. The confidence interval, first: if it overlaps with the previous value, caution is warranted. Then the descriptive categories — extremely low, borderline, average, high average, superior for a composite index; borderline and clinical range for a behavioral scale — because moving from one category to another is often more meaningful than the raw point difference.

Direction also matters. On a BDI-II or PHQ-9, a lower score means improvement; on an IQ index, it is the opposite. A tracking system that treats every increase as progress will be wrong half the time.

Brief scales as a supplement

Between full assessments, short screening tools — the GAD-7 for anxiety, the PHQ-9 for depression — provide a regular, low-burden reference point. Repeated at each appointment, they trace a trajectory that clinical impression alone does not always capture.

Feedback is part of the assessment

An assessment that is never fed back has not fulfilled its purpose. The feedback session — with the patient, with parents, sometimes with an adolescent separately — requires translating index scores into usable language without stripping them of meaning. "He picks up new concepts quickly when they are explained to him, but he needs more time than his peers to get things down in writing" says more than a table of standard scores.

What helps at that moment: having the full profile in front of you rather than just the global score, and being able to show change since the previous assessment when one exists. A family is far more likely to accept a placement recommendation when they can see what it is grounded in.

What leaves the file, and for whom

A psychological assessment is not meant for everyone. Conclusions and recommendations are shared with the referring clinician and, usually, the family; raw scores and detailed observations remain in the file. An assessment intended for personal use — working notes, a hypothesis still being developed — should not be visible to the rest of the team.

In practice

The allied health module in Hakim-DZ records each test with its version, date, index scores, and confidence interval, plots the indices from a single test on one graph with their descriptive categories displayed, and separates team-visible assessments from personal ones. The GAD-7 and PHQ-9 are built into the profession-specific fields.

See the Allied Health page, and the use case Tracking Rehabilitation Progress.