Somewhere in the recruitment process, a link arrives. Forty minutes later you have a colour, an animal, a four-letter code or a bar chart, and a stranger in human resources now holds a document that purports to describe you. It is worth knowing what that document can legitimately claim, because the honest answer varies enormously depending on which instrument produced it.

How assessment got into the office

Mass psychological testing arrived with the First World War, when armies needed to sort enormous intakes quickly and psychologists offered to help. The instrument usually credited as the first personality inventory came out of that effort: Robert Woodworth's Personal Data Sheet, a list of yes-or-no questions about nerves, sleep, fears and physical complaints, designed to flag recruits at risk of breaking down under shellfire. As Smithsonian Magazine recounts, it was finished too late to be used for its intended purpose โ€” but the format survived and multiplied.

Industry took the idea up between the wars and never really let go. By the middle of the century, testing had become a routine feature of white-collar hiring in the United States, sufficiently pervasive that it drew sustained criticism from writers examining corporate conformity. The pattern set then still holds: a measurement instrument built for one narrow purpose gets adopted for a much broader one, because it is cheap, it produces a tidy number, and it makes an uncomfortable judgement feel like an objective one.

What 'validity' means when a vendor says it

The word does a lot of concealment. In selection research, the relevant idea is criterion-related validity: the statistical relationship between how someone scores on the assessment and how well they subsequently do the job, measured against something outside the test. That last clause is where most workplace instruments quietly fail. A questionnaire can be internally consistent, reliable on retest, popular with candidates and completely uninformative about job performance, because none of those properties is the one being claimed.

Even genuinely validated methods are weaker than the marketing implies. A validity coefficient of 0.4 is towards the top of what selection science achieves, and it corresponds to explaining a modest fraction of the variation in performance. At the scale of a thousand hires, an advantage like that is real money. For the individual sitting the test, it means the prediction about them is wrong often enough that no sensible person would treat it as a verdict. Aggregate usefulness and individual accuracy are different animals, and conflating them is the single most common error in how these results get discussed.

The 2022 recalculation

The field itself had been overstating things, and said so. For decades, published estimates of how well various selection methods predict performance were adjusted upwards to correct for range restriction โ€” the statistical problem that you only observe job performance among people you actually hired. In 2022, Paul Sackett and colleagues published a re-examination in the Journal of Applied Psychology arguing that these corrections had been applied inappropriately and had systematically inflated the numbers.

The revised estimates rearranged the league table. Structured interviews came out at 0.42, down from a previously cited 0.48 but now at the top. Biodata, meaning empirically keyed questions about a person's actual history, rose slightly to 0.38. General cognitive ability tests fell hardest, from roughly 0.52 to 0.31, dislodging them from the top spot they had occupied in the literature for a generation. Integrity tests dropped from 0.42 to 0.31, and conscientiousness measures came in at 0.19.

Two things follow. The first is that structure, rather than the instrument, is doing much of the work โ€” a structured interview is not a different conversation, it is the same conversation with fixed questions, defined rating scales and more than one assessor, and that discipline is what raises it above the unstructured chat it replaced. The second is that the methods which perform best all share a family resemblance: they ask what the person has actually done, or ask them to do a sample of the work in question, rather than asking them to describe their own temperament.

Notice what is absent from these tables. Instruments that sort people into named categories or types generally do not appear, because they were not built to predict job performance and their more careful publishers say as much in the manual. There is also a measurement problem underneath: where a trait is genuinely continuous, chopping it at a midpoint discards information and produces classifications that can flip between sittings for anyone near the boundary. A tool can be a pleasant vocabulary for a team away-day and still be the wrong object entirely for deciding who gets a job.

The constraints nobody mentions in the away-day

Hiring assessment is not merely a scientific question, it is a regulated one. The US Equal Employment Opportunity Commission's guidance on employment tests and selection procedures lays out the framework plainly. A selection procedure that is neutral on its face but disproportionately excludes members of a protected group must be shown to be job-related and consistent with business necessity, and an employer may still be liable if a less discriminatory alternative would serve the same purpose. The Uniform Guidelines adopted in 1978 set out the validation evidence expected.

Disability law adds another layer. Assessments that screen out people with disabilities must clear the same job-relatedness bar, accommodations must be available in how the test is administered, and questions that stray towards medical inquiry are legally fraught โ€” which is awkward for personality inventories descended, as many are, from clinical screening instruments. Comparable pressure exists elsewhere: equality legislation and data-protection rules on automated decision-making constrain profiling in the UK and Europe along similar lines.

The professional standards are equally clear. The Principles for the Validation and Use of Personnel Selection Procedures, published by the Society for Industrial and Organizational Psychology, describe what evidence an employer should hold before using an assessment to make decisions about people. Plenty of organisations use tests they could not defend against that document for five minutes.

Why the weak instruments survive anyway

If the evidence is this available, the persistence needs explaining, and the explanation is mostly organisational rather than intellectual. Tests are cheap and scale effortlessly. They are pleasant to take, which matters more than it should, because candidates and staff rarely complain about being told something flattering. They convert an anxious managerial judgement into a document, which distributes responsibility if the hire goes wrong. Vendors sell certification courses that create internal advocates with a stake in continued use. And almost nobody closes the loop: it is rare for an employer to check, years later, whether the people the instrument favoured actually outperformed the people it did not.

There is also a use that is not really about prediction at all. A shared vocabulary for describing working styles genuinely helps teams talk about friction without it becoming personal, and a lot of workplace testing is performing that social function while wearing the costume of measurement. That is a defensible thing to want. The damage begins when the same output is carried into decisions about hiring, promotion or who gets moved off a project.

Reading your own result sensibly

A few questions separate a serious instrument from an expensive quiz. What outcome was it validated against, and can the employer produce that evidence? Was it scored against a norm group that resembles the people taking it? Does it report positions on continuous scales with some indication of measurement error, or does it hand you a label with no margins? Would you get the same result in three months, and does anyone know?

Then read the output for what it is: a summary of how you described yourself on one particular afternoon, filtered through the instrument's assumptions about what matters. Useful for a conversation with your manager about how you prefer to work. Reasonable evidence that you might dislike a role built entirely around things you find draining. Not a diagnosis, not a ceiling, and not a legitimate reason for anyone to conclude that a capable person cannot do the job.