← All work
Self-initiated studyHealthTechClinical-adjacentMeasurement design0 to 1

A self-screener rigorous enough to act on, honest enough not to diagnose

Neurodivergent Traits Screener, a self-initiated Product Office study, 2026

8
Trait dimensions screened in one tool
96
Items, engineered for reliability, not speed
60%
The flag line, labelled pragmatic, not clinical
0→1
Intent, architecture, scoring, content, build

The outcome first: a health tool that knows what it can't know

What I owned: The whole thing. I set the product intent, designed the two-path flow, wrote the scoring model, engineered a 96-item question bank for reliability, resolved every open decision in a written record, and specced the build. A self-directed study in how to make a measurement product honest.

I designed an intro-level self-screener that helps an adult work out which neurodivergent traits, if any, are worth a closer look, across eight dimensions from ADHD to dyspraxia to OCD. The hard part was never the questions. It was the restraint. A self-test in a health context can do real harm two ways: scare someone with a false positive, or falsely reassure them with a false negative. So the entire product is built to point, never to conclude. It routes people to the right specialised test and, past that, to a clinician. Every design call defends that line.

The problem I chose to solve: people screen themselves too narrowly

An adult who suspects they are neurodivergent hits a wall. The validated tools are long, siloed, and one per condition, so you have to already know what to test for. Most people don't. They land on a single label from a social feed ("I think it's ADHD") and screen only for that, missing the conditions that co-occur or the ordinary things (trauma, thyroid, chronic sleep loss, perimenopause) that mimic the same profile.

Nothing helps at the step before the specialised tests: figuring out where to look first. That gap, surfacing what you didn't think to consider without overclaiming, is the product.

The core decision: a router, not a verdict

Most self-tests online either hand you an authoritative-sounding result or keep you clicking. I designed this one to sit deliberately at the bottom of a three-layer funnel, and to say so out loud on the results page. That single framing decision drives everything downstream: what it may claim, how results are worded, where it sends you next.

Where the tool sits, and what it refuses to be

A three-layer funnel. This tool is layer one, on purpose. Each layer down is narrower, slower, and more certain. The screener's job is to get you to the right layer two. Layer 1 · Quick screen (this tool) Broad, intro-level. Points you toward which areas to look at. Not reliable enough to decide on alone. YOU ARE HERE Layer 2 · Specialised self-report ASRS, RAADS-R, OCI-R, etc. More reliable, but still self-report, not a diagnosis. Layer 3 · Clinical assessment A professional. Formal diagnosis via a GP referral. The results page teaches the user this map, so "elevated" reads as "take the next step", never "you have this".

From that spine, the flow splits into two entry paths, matched to how much the user already knows.

Door 1

Targeted screen

For people who already have a hunch. Pick the dimensions you want, answer the full 12-item screen for each. A "screen everything" option exists too, behind a length warning so nobody sleepwalks into 96 items.

"I know what I want to check"
Door 2, the recommended default

Broad scan

A 24-item triage across all dimensions flags what's worth a closer look. The user then confirms which to explore in full, and their triage answers carry forward so nothing is asked twice.

"I'm not sure what applies"

The entry screen, live: the two doors

The screen below is the real, clickable entry point, drawn straight from the build. The broad scan leads as the recommended default and the targeted path sits quieter beside it, so the person who doesn't know where to start is guided, not met with a wall of tests. The "not a diagnosis" line and the reminder about conditions that can mimic these profiles are on screen from the very first moment, not buried at the end.

traits-screener.app/start
Live prototypeThe landing and its two entry paths. Recommended broad scan first, targeted screen second. Open the targeted door to reveal the dimension picker with plain-language descriptions.

The decision that shaped the whole tool: twelve items per dimension, not five

Earlier versions used five items per dimension. I killed that. At five items, a single misread or ambiguous question moves the dimension score by twenty points, enough to flip someone between "low" and "elevated". A tool people might act on cannot swing on one question.

So I rebuilt each screen to twelve items, sized to clear the reliability floor a screener needs before anyone should trust it, and to properly cover the sub-facets of each dimension rather than rewording one idea. The cost is length. I paid for it deliberately, and then bought it back with the Door 2 triage and modular selection, so no one answers more than they chose to.

Why five items wasn't safe: one misread, and the band flips

The effect of one misread answer on the dimension score LOW · 0 to 39 MODERATE · 40 to 59 ELEVATED · 60 to 100 0 100 5 items one answer = 20 pts crosses two thresholds: low to elevated 12 items one answer = ~8 pts Illustrative. Reliability depends on item quality and correlation, not count alone, but more items means each answer moves the score less.

Designing against harm: restraint is the feature

The interesting product work was everywhere the honest choice fought the engaging one. I took the honest one each time, because in this domain trust compounds and overclaiming is a one-time con. A few of the calls:

Frequency scale, not agreement

"How often does this happen" instead of "do you agree this is you". It sidesteps self-image filtering and the acquiescence and central-tendency biases of agree/disagree scales, and it matches the real clinical tools, so scores translate cleanly if the user moves to the validated version.

A gentle flag, not a silent penalty

Reverse-scored attention checks catch straight-lining. When answers contradict, results say so kindly ("some answers seemed contradictory, consider retaking") rather than quietly degrading a score the user can't see.

Traits are not a diagnosis, built into results

A single impairment follow-up per elevated dimension. A 75% trait score with no daily impact means something different from the same score with severe impact, and the results show both, so "I have these traits" never reads as "I have this condition".

Name the mimics at the right moment

A differentials card (trauma, depression, anxiety, thyroid, sleep, perimenopause) shown on results, not before. That is the moment a user can actually act on "consider these alongside your score" when they take it to a GP.

Honesty as a feature

The 60% "elevated" line is a pragmatic cutoff, not a validated one, and the copy says exactly that wherever the number appears.

The items are written in the style of real instruments but are not empirically validated. The tool states this plainly. Owning the limit is what earns the right to be useful at all.

How I made the calls: decisions first, facts second, build last

I ran this like a product, not a side project. Ten open questions, from "should the broad scan auto-run or force a confirm" to "which extra dimensions earn inclusion", were each resolved in a written record: the decision, the options considered, and, crucially, what I rejected and why. That last part is the one people skip and later regret.

Before a single clinical claim could ship, I split the work by type. Design choices I could just make. Factual claims (follow-up test details, co-occurrence rates, diagnostic criteria) went into a separate verification track, sixteen research prompts sent out to be sourced before they were allowed near the copy. Knowing the difference between a decision and a fact is most of the discipline.

Add what you can operationalise honestly

Alexithymia earned a full dimension because a real adult instrument exists to route people to. Tics and auditory processing were deferred, not faked, until a valid self-screen is confirmed. Demand avoidance became a sub-flag inside autism, matching the clinical reality rather than inventing a dimension.

Give context its own shape

Giftedness doesn't screen cleanly as twelve items, so it isn't a dimension. It's a results-page note explaining why neurodivergence gets missed, alongside masking and late diagnosis. Different kinds of truth get different kinds of treatment.

Outcomes: what the study produced

8 in 1
Eight trait dimensions unified into one coherent screen with two flows, replacing a scatter of siloed single-condition tests
96 items
A question bank rebuilt from 5 to 12 items per dimension to clear the reliability floor a tool people act on actually needs
10 ADRs
Every product decision resolved and written down with its rejected alternatives, so the reasoning survives the session
16 checks
Clinical claims separated from design choices and sent for source verification before any of them reached the copy

What I'd take from this: the product was the discipline

Anyone can write a personality quiz. The work here was refusing to let it become one. Every strong decision came from the same place: a clear-eyed view of what the tool could honestly claim, and the discipline to design against the temptation to claim more. In a regulated, health-adjacent product, that restraint isn't a constraint on the work. It is the work.

Next case study
Reframe, turning an optician's instinct into an explainable product