Higher validity. Better talent decisions
The numbers behind the most predictive science on the market, and what that accuracy is worth in hires that work out and promotions that stick.
How well each assessment forecasts performance
Criterion-related validity is the correlation between an assessment score and an independent measure of how well someone actually performs at work, such as manager and peer ratings or objective outcomes like sales figures. Around 0.3 is widely considered good for a personality measure. Validities above 0.6 are extremely difficult to achieve for any job, whether using a single assessment or a combination.
| Assessment | Criterion validity |
|---|---|
| Wave® Professional Styles | 0.57 |
| Match 6.5 | 0.45 |
| Wave® Focus Styles | 0.44 |
| Situations SJTs | 0.28 |
| Wave® Professional Styles with Swift Aptitude | 0.61 |
| Match 6.5 with Swift Aptitude | 0.49 |
Wave® and Match figures based on overall competency score correlated with independently rated global work performance, adjusted for criterion unreliability (N=308). Performance was rated by managers, colleagues and others on a separate instrument, not by the person being assessed. Situations figure based on meta-analysis research study covering multiple Situations assessments created for different job roles and organizations. Overall Situations score correlated with soft and hard performance criteria, adjusted for criterion unreliability (k=9, N=1,900).
Our behavioral assessments demonstrate strong validity when used for hiring and development. Where cognitive ability is important for effective performance in a role, combining behavioral with aptitude assessment provides incremental validity and even stronger forecasting of performance.
What the difference between good and excellent is worth
A coefficient is abstract until you convert it into decisions. A serious selection error means appointing someone from the bottom 20% of performers when you set out to appoint from the top 20%. Here is how often that happens at each level of validity.
| Assessment validity | Approximate odds of a serious selection error |
|---|---|
| No assessment | 1 in 5 |
| ~0.30 | 1 in 10 |
| ~0.60 | 1 in 50 |
Moving from 0.30 to 0.60 does not halve your error rate. It cuts it by four fifths. That is the heart of the commercial argument for caring about validity.
Double the validity of your assessment and you can double the financial benefit it delivers. Same spend. Twice the gain.
How consistently the measurement holds
Reliability is how consistently an assessment measures. We report three kinds, because each one tests something different.
Alternate form asks whether two different versions of the questionnaire produce the same profile. Test-retest asks whether that profile still holds when the person completes it again later. Internal consistency asks how closely the questions within a single scale hang together.
| Measure | Wave® Professional Styles | Wave® Focus Styles |
|---|---|---|
| Alternate form | Median 0.87, range 0.78 to 0.93 (N=1,153) | Median 0.87, range 0.77 to 0.89 (N=504) |
| Test-retest | Median 0.74 at 18 month interval, range 0.58 to 0.85 (N=100) | Median 0.77 at six month interval, range 0.74 to 0.88 (N=214) |
| Internal consistency | Median 0.76, range 0.58 to 0.86 (N=1,153) | Median 0.74, range 0.59 to 0.86 (N=504) |
Professional Styles figures are across 36 dimensions, Focus Styles across 12 sections. No corrections applied to the reliability figures.
We lead on alternate form because it is the most meaningful. Internal consistency is the figure questionnaires often lead on, however a scale built from near-identical questions scores highly on it while measuring only a narrow concept. Our dimensions carry varied content by design, which is the right trade: breadth of measurement over a flattering number.
Fair by measurement, not by assertion
We publish mean score differences by age, gender and ethnicity for multiple samples, alongside differences by level of management responsibility and by region. Differences are reported as effect sizes, so they can be judged against a standard rather than described as small in our own words.
Differences that do not justify different treatment
Across age, gender and ethnicity, differences are generally negligible or small, and none give a reason to treat any group differently.
One norm group, one method
Separate norms by age, gender or ethnicity are not offered and not recommended. Consistency of method for a given role is what makes a decision hold up when someone asks about it later.
Adding Wave® can improve fairness
Adding Wave® to a selection process is unlikely to introduce adverse impact and may even reduce it.
What a score is compared against
By role level
Options include senior managers and executives, professionals and managers, individual contributors, graduates, apprentices, technical occupations, foundation level, mixed occupational, candidates completing in an additional language.
By geography
International, regional and country-level norms, updated regularly as data accumulates. The group used is recorded on every report, so the relevance of this comparison can be referenced.
Sized for accuracy, not for headlines
Standard error of the mean falls steeply up to around 500 cases and changes very little after that. Past that point, how representative a sample is matters far more than how large it is, which is the part a headline sample size never tells you.
Ask us for what you need
Technical manuals, validity and reliability summaries, fairness data and norm documentation go out on request, so the right material reaches the right reviewer and one of our psychologists is on hand to answer questions about it.
A psychologist reads every request. You will hear back from a person.
Go deeper
Research and resources
Whitepapers, webinars and podcasts from the people doing the research, including the leadership potential work and the wider validity program.
The science behind Wave®
How the model was built, why item selection was driven by prediction rather than theory, and what independent reviewers made of it.
Accreditation and training
Test User: Occupational, Ability and Personality, plus Wave® accreditation for the people who will interpret these numbers day to day.
Questions technical reviewers ask
Wave® Professional Styles reaches a criterion validity of 0.57 against independently rated global work performance, and 0.61 combined with our Swift Aptitude test (concurrent, corrected for criterion unreliability; N=308). Around 0.3 is good for a personality measure and validities above 0.6 are extremely difficult to achieve, whether using a single assessment or a combination.
Item choices from the development trial were cross-validated against external criterion on a separate standardization sample, then again using a research sample. Managers, colleagues and others rated participants' overall effectiveness and individual competencies on a separate instrument, and assessment scores were correlated with those ratings.
Alternate form reliability has a median of 0.87 across the 36 dimensions of Wave® Professional Styles, ranging from 0.78 to 0.93, with no corrections applied. Test-retest at 18 months has a median of 0.74 (N=100). Internal consistency ranges from 0.58 to 0.86, with a median of 0.76, and is deliberately not maximized, because a scale of near-identical questions scores highly on it while measuring a narrow concept.
Group differences by age, gender and ethnicity are generally negligible or small across multiple samples. Adding Wave® to a process is unlikely to introduce adverse impact and may even reduce it.
Yes. A local validation study correlates assessment scores against your own performance data, usually concurrently on current employees. Our psychologists design and run these with clients. Ask us what would be involved for your roles.
Usually the one matching the role level and the most applicable region. Depending on the assessment, we maintain norms for senior managers and executives, professionals and managers, individual contributors, graduates, apprentices, technical occupations, foundation level and mixed occupational groups, across international, regional and country samples. Whichever is used is recorded on the report.
Get the evidence for your review
Tender and RFP responses, supplier due diligence questionnaires, equality impact assessments, works council consultations, internal business cases. Each asks for this evidence in a different shape. Tell us which one you are completing and we will send what fits.