THE
PERFORMANCE LENS
Updated July 2026
Methodology & Validation

The data behind
the profile.

The Performance Lens measures how people behave under pressure, and produces data you can trust.

01
StandardisedEveryone runs the same simulations, scored the same way. Differences reflect behaviour, not the test.
02
BehaviouralDrawn from what people did under real pressure, not how they describe themselves.
03
Multi-raterEach score blends a person's own view with the read of the people who worked alongside them.
04
Research-anchoredBuilt on a validated assessment format and grounded in established and current behavioural research.
05
ObjectiveScoring is automated and rule-based. The same responses always produce the same result.
The question this answers

Can the data be trusted?

When a leader commissions an assessment like this, the serious first question is rarely what the report looks like. It is whether the data underneath it can be trusted: whether it is accurate, whether it is fair, and whether it means anything beyond an opinion formed on the day.

This page answers that question directly. It sets out what The Performance Lens measures, how the underlying data is produced and scored, and why you can rely on the result.

The Performance Lens is a diagnostic system for developing people and teams. Rather than asking people to describe themselves on a survey, it puts every participant through structured simulations, observes how they behave when the exercise puts them under pressure, and scores that behaviour against research-backed patterns of what effective performance looks like. Every workshop we run, current and future, uses this same method; only the specific behaviours it measures change. The output is an individual behavioural profile and a team-level view for the leader.

01 · What we measure

The behaviours that decide performance

Performance is not a single quality. Every Performance Lens assessment measures the handful of behaviours that most determine success in the capability that workshop develops, whether that is working as a team, leading others, or handling a specific challenge. Each behaviour is defined in observable terms, so a score points to something a person did, and can develop, rather than to a fixed trait or personality label.

The dimensions are specific to each workshop. What does not change is the method for choosing and defining them.

Team Effectiveness
CommunicationDecision-MakingCollaboration
Leadership
Leading with IntentionNavigating ChangeManaging Conflict

Each workshop, current and future, brings its own dimensions, chosen and defined the same way. The method is modular, so we can also build a bespoke workshop for a client, combining skills from across our library into a custom assessment scored to the same standard.

Three lenses on every person

Each question reads a person through one of three lenses, so the profile reflects the whole of how someone operates, not a single angle.

Strategic
How they think

How a person reads a situation, plans, weighs options, and judges their own confidence.

Emotional
How they feel

How a person manages themselves when a task is ambiguous, constrained, or not going to plan.

Behavioural
How they act

What a person does: their first move, and how they carry it out with others.

02 · How the data is generated

Pressure reveals defaults

People can describe themselves generously on a questionnaire. They behave as they are when a task constrains them and the clock is running.

Every participant completes the same set of structured simulations, designed so the behaviours we want to measure are forced to the surface: information is deliberately limited, roles are assigned and then disrupted, and the exercise cannot be completed well without communicating, deciding, and coordinating in real time. What a person does inside that controlled setting is the raw material of the assessment.

Everyone runs the identical instrument

Because the simulations are the same for every participant and every cohort, results are directly comparable, person to person and group to group. A difference between two people reflects a difference in behaviour, not in the exercise. Standardisation is what allows a score to carry a stable meaning at all.

Reflection is anchored to a real, recent experience

Participants respond to what they have just done, moments after doing it. Retrospective report about a concrete, recent experience is far more reliable than a hypothetical "what would you do if" question, which measures aspiration rather than behaviour. The exercise creates the evidence; the questions ask the person to account for it.

Why this is stronger than a survey

A personality survey records how a person wishes to be seen. A behavioural simulation records how a person operated when a real task demanded it. The gap between those two things is exactly what a team leader needs to see, and only the second approach can show it.

03 · Two views of every person

No one is scored on their own opinion alone

The single most important safeguard in the method is that each person's scores blend two independent sources.

The first is the person's own account of how they approached the task. The second is the read of the people who worked alongside them and watched them operate. Both views are captured on the same dimensions and the same scale, and the final score combines them. This is a standard move in team and 360-style assessment, and it directly reduces the two well-known weaknesses of pure self-report: social desirability (the pull to answer as one would like to be seen) and simple inaccuracy in how people perceive their own behaviour.

Agreement and divergence both carry signal

Where a person's own view and their colleagues' view agree, confidence in the score is high. Where the two diverge, the assessment has surfaced a genuine gap between how a person operates and how they land with the people around them. Either way, the combined score is more robust than either view alone.

These colleagues report on observable behaviour, what a person did, rather than a verdict on whether they were "good." That keeps the input honest and inside a behavioural frame. The peer contribution is blended quietly into the dimension scores. Nothing on an individual's profile announces that peers rated them; the profile a participant receives is private, individual, and framed for development. The method gains the accuracy of multiple raters without putting anyone under the discomfort of visible peer judgement.

04 · How responses become scores

Simple to state, hard to fault

The scoring model runs the same way for every participant, with no discretion applied to the result.

One response, one dimension

Every response maps to exactly one of the dimensions, based on what the response reveals about behaviour.

A research-anchored behavioural scale

Within each dimension, responses are weighted on a graded scale. Responses that more fully reflect effective behaviour in that specific situation carry more weight; responses that reflect a less effective approach carry less. Every option describes a real professional behaviour, so this is a scale of effectiveness in context, never a judgement of the person.

Normalised to a common scale

Raw scores are converted to a 0 to 100 scale for each dimension, adjusted for the exact set of questions each participant answered, so people who held different roles remain directly comparable. The self and peer components are each placed on this common scale and then combined.

Automated end to end

Collection, scoring, normalisation, and profile generation run through an automated, deterministic pipeline: identical inputs produce identical outputs, every time. No part of the score depends on who facilitated the session. This is the same discipline modern judgement research calls for: removing the inconsistency, or "noise," that creeps in when humans rate each other by hand (Kahneman, Sibony & Sunstein, 2021).

What this buys you

Objectivity and consistency. Two people who behave the same way receive the same score. The same person assessed by a different facilitator receives the same score. The number is a property of the behaviour, not of the room.

05 · The research foundation

Grounded in the science, kept current

Is this kind of test trustworthy?

Situational judgement tests are one of the most heavily validated assessment formats of the last two decades, with meta-analytic evidence on both their validity and their reliability (Christian et al., 2010; Webster et al., 2020). The Performance Lens is built as one. So the first question a sceptic asks, whether a test of this kind can be trusted, is already answered by decades of published research, before we say anything about our own instrument.

On top of that validated format, the dimensions in each workshop, and the way we rank behaviour within them, are grounded in established research and kept current with where that research has moved. The constructs are recognised outside the product, which is what lets the profile claim it measures things known to matter. The dimensions for each workshop are shown below, with the research each one anchors to.

Team Effectiveness

DimensionAnchoring concepts and sources
CommunicationClosed-loop communication and grounding in dialogue (McIntyre & Salas, 1995; Clark & Brennan, 1991), within the current synthesis of teamwork competencies (Salas, Reyes & McDaniel, 2018).
Decision-MakingNaturalistic and recognition-primed decision making (Klein, 1998); confidence calibration (Moore & Healy, 2008); learning through after-event review (Ellis & Davidi, 2005); and the case for removing human inconsistency from judgement (Kahneman, Sibony & Sunstein, 2021).
CollaborationShared mental models and backup behaviour (Cannon-Bowers, Salas & Converse, 1993; Salas, Sims & Burke, 2005); and psychological safety, from the foundational study to its modern form (Edmondson, 1999; 2019), reinforced by Google's Project Aristotle (Rozovsky, 2015).

The Leadership Workshop

DimensionAnchoring concepts and sources
Leading with IntentionTransformational and authentic leadership (Bass & Riggio, 2006; Walumbwa et al., 2008), within the modern leadership-development evidence base (Day et al., 2014).
Navigating ChangeEstablished change models (Lewin, 1947; Kotter, 1996) and adaptive leadership (Heifetz, Grashow & Linsky, 2009), with the modern review of how people respond to change (Oreg, Vakola & Armenakis, 2011).
Managing ConflictConflict-handling styles (Thomas & Kilmann, 1974; Rahim, 2002), the evidence on task versus relationship conflict (De Dreu & Weingart, 2003), and principled negotiation (Fisher & Ury, 1981).
Navigating Failure is in development. Its dimensions and research anchors will be published here when the workshop launches, drawing on the research on resilience, learning from failure, and recovery.

These sources anchor the format, the constructs, and the behavioural rankings. They describe the theoretical and evidentiary foundation; they are not a claim that this specific instrument has been independently validated against each source. Our own validation and calibration are continuous, and are described in the next section.

References

06 · How we keep it current

A method that improves with every cohort

An assessment earns trust by staying current and improving in the open. Our review runs on three cadences.

We would rather state plainly what the method already does well, and keep sharpening the rest, than overclaim on either.

07 · Designed to be trusted

Clean inputs, fair to everyone

The scoring is invisible, so it cannot be gamed

Participants never see how responses are weighted, and every answer option describes a genuine, credible professional behaviour. No option is written to be obviously the "right" one, so a participant cannot reverse-engineer the scoring and select their way to a flattering result. This is what keeps responses honest, and it is why the multi-rater blend can do its job.

Fair across language and context

The assessment is written in plain English suited to a workplace where many participants speak English as an additional language. There is no time pressure, so a score reflects considered behaviour rather than reading speed, and each person responds independently, so answers are not pulled toward a group consensus.

Framed for development, and kept private

Every individual profile is private to that participant and written in constructive, growth-oriented language. Alongside the scored dimensions, a separate layer captures each person's working-style preferences. That layer is non-evaluative by design: preferences are described, never scored, because a preference is not a strength or a weakness.

08 · The evidence so far

What people in the room are saying

The early signal from the field is strong on the two things that matter most for a diagnostic: participants find it credible, and they find it useful the next day.

4.7/5
Average workshop rating
150+
Professionals assessed
91%
Want a future session
92%
Rated the content immediately applicable
"Learning by doing is the best way in developing a new muscle memory. The way of engagement & gamification makes the workshop fun! Most importantly, the self awareness created & also the observation of others brings insights."

Christine Ng, Founder, Why Ventures

Trusted by

Figures reflect assessed participants to date and update as the dataset grows.

09 · One method, one standard

The same engine across the family

Everything described here is a single measurement engine: standardised behavioural simulation, a self and peer blend, a research-anchored scoring scale, and an automated, deterministic pipeline. That engine is the standard every method in the family is built to. Only the dimensions, the exercises, and the profile language change; the way behaviour becomes trustworthy data does not.

Team Effectiveness Leadership Navigating Failure · in development

The Team Effectiveness and Leadership assessments both run on this methodology. Team Effectiveness is the worked example throughout this page; the Leadership assessment applies the same engine to leading with intention, navigating change, and managing conflict. A further method, on navigating failure, is in development. Because the method is consistent, a partner who is satisfied with the data from one is looking at the same evidentiary standard across all of them.

See it. Fix it. Prove it.

This page sets out the method at the level a serious buyer needs to make a decision. The instrument is proprietary, as with any established assessment; we're glad to walk your team through the method and answer any question.

Talk to us about the methodology