Skip to content
Search
The evidence

Higher validity. Better talent decisions

The numbers behind the most predictive science on the market, and what that accuracy is worth in hires that work out and promotions that stick.

Validity

How well each assessment forecasts performance

Criterion-related validity is the correlation between an assessment score and an independent measure of how well someone actually performs at work, such as manager and peer ratings or objective outcomes like sales figures. Around 0.3 is widely considered good for a personality measure. Validities above 0.6 are extremely difficult to achieve for any job, whether using a single assessment or a combination.

AssessmentCriterion validity
Wave® Professional Styles0.57
Match 6.50.45
Wave® Focus Styles0.44
Situations SJTs0.28
Wave® Professional Styles with Swift Aptitude0.61
Match 6.5 with Swift Aptitude0.49

Wave® and Match figures based on overall competency score correlated with independently rated global work performance, adjusted for criterion unreliability (N=308). Performance was rated by managers, colleagues and others on a separate instrument, not by the person being assessed. Situations figure based on meta-analysis research study covering multiple Situations assessments created for different job roles and organizations. Overall Situations score correlated with soft and hard performance criteria, adjusted for criterion unreliability (k=9, N=1,900).

Our behavioral assessments demonstrate strong validity when used for hiring and development. Where cognitive ability is important for effective performance in a role, combining behavioral with aptitude assessment provides incremental validity and even stronger forecasting of performance.

See how the assessments differ

What it is worth

What the difference between good and excellent is worth

A coefficient is abstract until you convert it into decisions. A serious selection error means appointing someone from the bottom 20% of performers when you set out to appoint from the top 20%. Here is how often that happens at each level of validity.

Assessment validityApproximate odds of a serious selection error
No assessment1 in 5
~0.301 in 10
~0.601 in 50

Moving from 0.30 to 0.60 does not halve your error rate. It cuts it by four fifths. That is the heart of the commercial argument for caring about validity.

Double the validity of your assessment and you can double the financial benefit it delivers. Same spend. Twice the gain.

See the return this produces

Reliability

How consistently the measurement holds

Reliability is how consistently an assessment measures. We report three kinds, because each one tests something different.

Alternate form asks whether two different versions of the questionnaire produce the same profile. Test-retest asks whether that profile still holds when the person completes it again later. Internal consistency asks how closely the questions within a single scale hang together.

MeasureWave® Professional StylesWave® Focus Styles
Alternate formMedian 0.87, range 0.78 to 0.93 (N=1,153)Median 0.87, range 0.77 to 0.89 (N=504)
Test-retestMedian 0.74 at 18 month interval, range 0.58 to 0.85 (N=100)Median 0.77 at six month interval, range 0.74 to 0.88 (N=214)
Internal consistencyMedian 0.76, range 0.58 to 0.86 (N=1,153)Median 0.74, range 0.59 to 0.86 (N=504)

Professional Styles figures are across 36 dimensions, Focus Styles across 12 sections. No corrections applied to the reliability figures.

We lead on alternate form because it is the most meaningful. Internal consistency is the figure questionnaires often lead on, however a scale built from near-identical questions scores highly on it while measuring only a narrow concept. Our dimensions carry varied content by design, which is the right trade: breadth of measurement over a flattering number.

Fairness

Fair by measurement, not by assertion

We publish mean score differences by age, gender and ethnicity for multiple samples, alongside differences by level of management responsibility and by region. Differences are reported as effect sizes, so they can be judged against a standard rather than described as small in our own words.

Differences that do not justify different treatment

Across age, gender and ethnicity, differences are generally negligible or small, and none give a reason to treat any group differently.

One norm group, one method

Separate norms by age, gender or ethnicity are not offered and not recommended. Consistency of method for a given role is what makes a decision hold up when someone asks about it later.

Adding Wave® can improve fairness

Adding Wave® to a selection process is unlikely to introduce adverse impact and may even reduce it.

Norms

What a score is compared against

By role level

Options include senior managers and executives, professionals and managers, individual contributors, graduates, apprentices, technical occupations, foundation level, mixed occupational, candidates completing in an additional language.

By geography

International, regional and country-level norms, updated regularly as data accumulates. The group used is recorded on every report, so the relevance of this comparison can be referenced.

Sized for accuracy, not for headlines

Standard error of the mean falls steeply up to around 500 cases and changes very little after that. Past that point, how representative a sample is matters far more than how large it is, which is the part a headline sample size never tells you.

Technical documentation

Ask us for what you need

Technical manuals, validity and reliability summaries, fairness data and norm documentation go out on request, so the right material reaches the right reviewer and one of our psychologists is on hand to answer questions about it.

A psychologist reads every request. You will hear back from a person.

Go deeper

Research and resources

Whitepapers, webinars and podcasts from the people doing the research, including the leadership potential work and the wider validity program.

The science behind Wave®

How the model was built, why item selection was driven by prediction rather than theory, and what independent reviewers made of it.

Accreditation and training

Test User: Occupational, Ability and Personality, plus Wave® accreditation for the people who will interpret these numbers day to day.

Questions technical reviewers ask

Wave® Professional Styles reaches a criterion validity of 0.57 against independently rated global work performance, and 0.61 combined with our Swift Aptitude test (concurrent, corrected for criterion unreliability; N=308). Around 0.3 is good for a personality measure and validities above 0.6 are extremely difficult to achieve, whether using a single assessment or a combination.

Get the evidence for your review

Tender and RFP responses, supplier due diligence questionnaires, equality impact assessments, works council consultations, internal business cases. Each asks for this evidence in a different shape. Tell us which one you are completing and we will send what fits.