Administered, Not Observed: What Remote Testing Misses

July 30, 2026

What remote testing removes from the clinical record, and why the score report never shows what is missing

A remote battery does not return a blank where the behavioral observation should be. It returns a clean, complete, professional-looking profile with a hole in the middle that nobody can see.

That is the problem in one sentence. Not that the scores are wrong. The scores are usually computed correctly, faster than most of us managed by hand, and nobody is nostalgic for the arithmetic. The problem is that the information required to decide what those scores mean was never collected, and its absence leaves no trace in the output. The report looks finished. Nothing on the page indicates which parts of the examination were conducted blind.

This is a narrow claim and worth stating narrowly. The room does not make an examiner right. It gives them more chances to notice that they might be wrong.

The examiner is an instrument

It is the only instrument in the battery with no normative table, no manual, no license fee, and no line item on the invoice. It is also the one most responsible for deciding whether the other instruments produced anything usable. It is the only instrument you already own, which may be part of why it is the easiest one to stop calibrating.

What that instrument records has no field on the score sheet. Whether the patient arrived dehydrated. Whether there was a large iced coffee on the desk and no water for three hours. Whether the hands were tremulous at intake and steadier by hour two. Whether the pause before the sixth digit was retrieval effort, distraction, or a wave of something physical passing through and receding. Whether the patient reached for the desk. Whether breathing changed. Whether a subtest failure was preceded by ninety seconds of visible autonomic distress that resolved before the next task began.

A camera captures a face and roughly two feet of torso, degraded, latency-delayed, and framed by the patient. The patient chooses the frame, and nobody has ever chosen a frame that included the desk. It does not capture the room, or what is sitting on that desk just outside it, and it cannot reliably separate a patient who is disengaged from a patient who is unwell. Those two produce similar numbers.

The examiner is a fallible instrument too.1 But fallible information is different in kind from information that was never available to be considered.

Remote testing is also not one condition. An examiner using a tablet in the room has not surrendered the room, and what follows concerns the unproctored home version.2

What remote and tablet administration genuinely gains

Remote administration reaches people who would otherwise go unassessed: rural patients, medically fragile patients, patients who cannot drive, patients whose nearest neuropsychologist is four hours away and booked into next year. Digital stimulus presentation is more uniform than a worn easel handled by six different examiners, which is a low bar and worth clearing regardless. Timing is captured mechanically rather than by thumb. Automated scoring reduces opportunities for transcription and conversion error. Item-level response data can be captured in ways paper never allowed.

And scheduling throughput improves, sometimes dramatically. That last one is worth separating from the others. It is the gain most likely to drive adoption and least likely to appear in the stated rationale.

The equivalence evidence is stronger than critics of remote testing tend to admit, and digit span in particular has among the better telehealth support of any measure in common use.3 That is worth stating plainly, and it is not the point. Equivalence of scores is not equivalence of clinical information.

The question is not whether those gains are real. It is whether they answer the particular referral question without removing information needed to interpret the result.

Why the loss is easy to normalize

Every psychologist doing assessment work depends on a very small number of commercial publishers. That is not a scandal. It is a structural condition of the field, and it explains why a missing observation is so easy to overlook.

Publishers are businesses, which they would not dispute. Revision cycles, platform migrations, licensing models, and the per-administration pricing of digital delivery are decisions made by companies with revenue targets, and those decisions sit alongside psychometric considerations rather than strictly downstream of them. Some of the equivalence evidence supporting digital and remote administration has been produced or sponsored by the same entities selling the delivery platform.

The technical detail a clinician needs in order to reason carefully about an instrument is available. It is available for purchase.

None of that is an accusation of bad faith. It is a description of incentives, and the reason to state it is simpler than any complaint about cost. Responsibility for the interpretation does not transfer to the publisher. It stays with the person who signs the report. A platform that computes flawlessly and prints instantly has not assumed any part of that burden.

The second loss

The first loss concerns what was happening around the score. The second concerns whether the clinician still knows what the resulting number can support.

Platform-based administration lets a clinician deliver a test competently without ever learning why it is built the way it is. Prompts appear, responses are entered, scores populate. The work looks identical from the outside, and this is precisely the difficulty: it also looks identical from the inside.

What is missing is the accumulated understanding of how a measure came to exist, what it was derived from, what it has been validated against, and what its numbers stop meaning when the conditions shift.

Digit span had a second job

The case that follows turns on a discrepancy within digit span. The structural history explains why a familiar-looking number can invite an outdated interpretation.

In adult neuropsychology, digit span has long served two functions. It measures auditory attention and working memory, which is the function it advertises. Scores derived from it, including reliable digit span, its revised variants, and the age-corrected scaled score, have also been studied as embedded performance validity indicators.4 Many clinicians learned that second use through supervision and the validity literature, not from the administration manual.

The WAIS-5 changed the structure of the task. It separated the previously combined components, reports them independently, and added Running Digits as a new updating measure.5

Whatever the psychometric rationale, the practical consequence is not subtle. Cutoffs do not carry onto a restructured test by assumption.6 The ones in circulation were derived on the prior edition.

A clinician who knows that history treats the gap as a live problem and compensates elsewhere in the battery. A clinician who knows only what the platform displays sees a scaled score, treats it as continuous with everything they were taught, and interprets with confidence. The platform will not flag the discontinuity. The platform will always print a number. That is, to be fair to it, exactly what it was built to do.

Samantha, below, passed her validity measures. The history still matters, because it shows how easily a familiar score can outlive the assumptions built around it.

The danger is greatest when administration, scoring, and interpretation are separated so completely that nobody remains close enough to the encounter to notice what the numbers cannot explain. Clinicians formed inside that arrangement are not careless. They were never given the conditions in which that knowledge forms.

A case, entirely fictional, assembled from a familiar shape

Samantha is twenty-four, a college senior preparing for the MCAT, and by every external measure high-functioning. She has been taking a prescribed stimulant intermittently since her late teens, written by the family internist who has known her since childhood, filled every few weeks and used almost exclusively for study blocks and exams. Lately the blocks are longer and the intervals shorter. When the medication runs late and sleep will not come, she layers on an antihistamine, or melatonin, or both.

She develops a fog she cannot describe well. Not sleepiness, not exactly confusion. Something in between, and frightening because it is unfamiliar. Her physician, working with what she reported, adjusts the dose upward. She feels better for a day. On the third day she has headache, nausea, fatigue, irritability, cramping, an off-schedule cycle in a body that had previously kept time like a train timetable, a racing heart, and intermittent vertigo.

The workup is unremarkable. Vitals fine, exam fine, labs showing low vitamin D and a low-normal potassium, nothing anyone would act on. She is told it is stress.

In the interval she does what a frightened, intelligent, well-resourced patient does at two in the morning. She searches. Then she asks a chatbot, and then she asks it again with the question shaped a little differently. The answers keep arriving fluent and complete and calibrated to the fear she brought. She refines the question and the answers improve, in the sense that they become more certain.

She remembers an unremarkable knock to the head during a soccer game months earlier, a knock that left no mark and no pain and that she had not thought about since, and it acquires new significance somewhere around the fourth search. She cycles through tumor. She cycles through cancer. At some point she thinks: am I just a hypochondriac? And then feels worse for having thought it, because now the symptoms have a second explanation and both of them are her fault. She presents to an emergency department in acute panic, is seen, is calmed, and goes home. The symptoms return.

So she does what a motivated patient with resources does next. She finds an evaluation that can see her within the week and completes it remotely.

The profile comes back with very high verbal comprehension, very high untimed reasoning across verbal and visual domains, mildly low processing speed, low working memory, and simple span near the 2nd percentile, against notably better performance on the updating task. Sustained attention measures fall below expectation. Standalone performance validity measures are administered and passed, which makes a broad invalid-performance explanation less likely. The examiner is experienced and conscientious, and reads the working memory and speed findings as likely state-related, attributes the picture to stress with an attention condition to be ruled out, notes that she was medicated at the time of testing, and recommends extended time.

The report called the findings state-related without establishing whether that state was typical of her, and then used them to support an accommodation.

She arranges a full-length practice administration with extended time. Her performance remains poor.

Some weeks later, after a vacation, real sleep, and ordinary hydration, she sits the MCAT and scores well. She concludes she had simply been tired.

What the profile could not tell them

The evaluation had captured performance during an unstable physical state and formatted it as a stable cognitive profile.

She was sleep-deprived, acutely dysregulated on a recently increased stimulant dose whose timing nobody had mapped against her testing window, over-caffeinated, and underhydrated. She was also medicated at the time of testing. Her lowest scores fell in the domains the medication was meant to support. That is not an explanation. It is another reason to pause.

She later described vertigo arriving and receding across the session. Nobody had documented when. The discrepancy between her simple span and her updating performance should have prompted questions about whether her state was fluctuating. It was read as a trait instead.

That pattern was not an explanation. It was the reason to stop. A profile in which an active updating task substantially outperforms simple span cannot be treated as a stable working memory result until moment-to-moment state has been examined.7 It should have triggered an immediate pause, a state check, and a decision about whether the examination could continue. Set against reasoning scores in the very high range, a 2nd percentile span is not a data point. It is an alarm.

In my own practice, a discrepancy like that earns a pause before it earns an interpretation. That means checking symptoms, sleep, hydration, medication timing, and whether the patient appears able to continue, before deciding what the score can support.

A score can be validly produced and still be clinically unstable, because the condition producing it was never characterized. The later MCAT result shows one thing: the original performance was not stable enough to carry the weight the report placed on it. It does not identify which factor mattered, and it settles nothing about the accommodation.

In a room, those clues are available. You see the cup. You see no water. You see the flush, the hand on the desk, the two-second pause that is not cognitive. You stop, you offer water, you take a break, you note the time and the last dose, you re-administer or you document that you could not. Over three or four hours with breaks and unstructured talk, an undisclosed medication history has many more openings to surface than it does in a compressed remote block where the patient is performing competence into a webcam.

Those clues are partial, and someone still has to read them. What the room buys is more chances to notice that the conditions may not support an interpretation at all.

The format made the missing information harder to observe and easier to mistake for complete data. It did not make the information unobtainable. Hydration, sleep, last dose, caffeine load, and symptom fluctuation can all be asked about directly from a screen. Those questions narrow the gap. They cannot recover what the patient did not notice, did not know to report, or could not place in time.

The part that lasts

The report was not merely unhelpful. It was actively costly.

She was told, in effect, that her body was fine and her mind was strained. She had come in with an unstable physical state and left with a rule-out and an accommodation aimed at a stable cognitive weakness. When she later recovered by drinking water and sleeping, she did not conclude that the evaluation had missed something. She concluded that she had been dramatic.

She learned to override her own signals. She learned to push through, to distrust the body’s report, to treat unexplained symptoms as evidence of her own excess. She is going into medicine, where she will be asked to evaluate other people’s unexplained symptoms, and people tend to extend to patients the discount they were taught to apply to themselves.

An assessment that produces nothing has a cost. A confident, wrong, professionally formatted answer costs more, and the cost compounds.

The questions this leaves

Not whether teleneuropsychology is valid. That framing invites a defensive answer instead of a clinical decision. Not whether digital platforms are good or bad. They are tools, sold by companies, with the strengths and blind spots that implies.

Two questions are worth carrying into the next scheduling decision.

What must be observed in this particular case in order to interpret the results, and can I observe it from where I am sitting?

Do I understand the origin and the limits of the number I am about to interpret?

The report is a clinical argument. When the author did not adequately observe the conditions under which the data were produced, the result is not a weaker version of the same document. It is a different kind of document, and the reader has no way to know the difference.

A test can be administered without being observed. A report should never pretend those are the same thing.

Notes

1.  Bicego, Vogel, and Kendra (2026), Archives of Clinical Neuropsychology, a qualitative content analysis of clinical neuropsychologists’ experience of behavioral observation, describes its role in contextualizing test findings alongside its dependence on trained judgment and the difficulty of formalizing it. It is a small interview study drawn from dementia practice and is offered here as conceptual support rather than as evidence of reliability across settings. In-person testing has departures of its own from the conditions the norms were built on: noisy offices, bedsides, interpreters, and rooms nothing like the standardization sample. Every one of those requires interpretation and documentation. On the routine gap between in-person testing conditions and standardization conditions, see the policy review in note 2. https://doi.org/10.1093/arclin/acag019

2.  Remote administration is not one condition. Sperling and colleagues (2024), Archives of Clinical Neuropsychology 39(2), 227–248, report strong foundational evidence for the acceptability, feasibility, and reliability of tele-neuropsychological testing using particular tests, under certain conditions, in specific settings, and with specific patient populations, while noting a dearth of research on in-home testing specifically and relatively few randomized studies of its reliability and validity. Marra and colleagues (2020), The Clinical Neuropsychologist 34, 1411–1452, reach compatible conclusions. Tablet administration with the clinician present preserves in-room observation. https://doi.org/10.1093/arclin/acad066  https://doi.org/10.1080/13854046.2020.1769192

3.  Brearly and colleagues (2017), Neuropsychology Review 27(2), 174–186, systematically reviewed twelve counterbalanced crossover studies of videoconference against on-site administration in adults with mean ages from 34 to 88. Heterogeneity precluded interpretation of a pooled summary effect. Across 497 participants, test-specific analyses found verbally mediated tasks, including digit span, verbal fluency, and list learning, unaffected by videoconference administration; the digit span analysis drew on five studies and 359 participants and produced a small nonsignificant effect. Boston Naming Test scores fell about a tenth of a standard deviation below on-site scores, as did untimed tasks and those allowing repetition. Heterogeneous data precluded meaningful interpretation of motor-dependent tasks, and studies with older participants and slower connections were more variable. The authors supported videoconference administration of verbally mediated tasks by qualified professionals using existing norms. These were controlled crossover designs rather than unproctored home administration, and no included sample was drawn from adults in their twenties. The review was supported by the Department of Veterans Affairs rather than by a test-platform vendor. Vendor-produced equivalence work exists separately; see the Q-interactive technical report series. https://doi.org/10.1007/s11065-017-9349-1

4.  Reliable digit span and related embedded indices are not universal detectors. Their operating characteristics vary by cutoff and by population, and sensitivity is often limited. Loring and colleagues (2016), Archives of Clinical Neuropsychology, found failure rates of 34 percent in early Alzheimer disease and 14 percent in amnestic mild cognitive impairment against 8 percent in controls at the commonly used cutoff of 7 or lower. Maiman and colleagues (2019), Archives of Clinical Neuropsychology, found conventional cutoffs produced inadequate specificity in an adult epilepsy sample. The authors concluded that cutoffs derived from mixed clinical groups produce unacceptably high false positive rates in those populations, and that combining embedded indicators lowers them. These indices belong inside a multimethod performance validity assessment rather than serving as a substitute for one. https://doi.org/10.1093/arclin/acw014  https://doi.org/10.1093/arclin/acy027

5.  The separated components are documented in Pearson’s WAIS-5 materials and sample reports, and all twenty WAIS-5 subtests, including Digits Forward, Digits Backward, Digit Sequencing, and Running Digits, are enumerated in Canivez, Watkins, McGill, and Dombrowski, construct validity of the WAIS-5, Assessment, advance online publication. Digit Sequencing and Running Digits are designated primary working memory subtests and Digits Forward and Digits Backward secondary, per the publisher’s comparison materials. https://doi.org/10.1177/10731911251412219  https://www.pearsonassessments.com/content/dam/school/global/clinical/us/assets/wais-5/wais-5-comparison-flyer.pdf

6.  Traditional reliable digit span is computed from the longest forward and backward spans passed on both trials, so with both subtests still present the computation remains performable. The published cutoff and classification-accuracy literature is anchored to WAIS-IV administration, and the WAIS-5 changed the administration sequence, the norms, and the subtest structure the index sat inside. Whether the established operating characteristics hold under the new structure is a question for validation rather than assumption. The discontinuity is sharper for the age-corrected scaled score and for revised variants incorporating sequencing, since the combined subtest they were computed from no longer exists. https://doi.org/10.1093/arclin/acw014  https://www.pearsonassessments.com/content/dam/school/global/clinical/us/assets/wais-5/wais-5-comparison-flyer.pdf

7.  Subtest-level discrepancy interpretation has well-known limits. Scatter is common, and difference scores are generally less reliable than the scores they derive from, so an unusual discrepancy calls for consideration of reliability and base rates rather than direct causal interpretation. Reynolds (1997), Archives of Clinical Neuropsychology, working with a large child and adolescent sample, argued that forward and backward span represent distinct processes and should not be combined for clinical interpretation. Gignac, Reynolds, and Kovacs (2019), Assessment, using data modeled on the WAIS-IV normative sample, estimated the model-based reliability of the combined Digit Span score at .74, well below its published stratified alpha, and cautioned against interpreting that composite. Nothing in the present case is offered as a performance validity finding; standalone measures were passed. https://doi.org/10.1093/arclin/12.1.29  https://doi.org/10.1177/1073191117748396

You Might Also Like