Pentest Team Performance Framework

A practical framework for measuring penetration testing delivery flow, quality, client friction, senior dependency, and consultant development without relying on utilisation alone.

Pentest team performance means moving work from a clear, agreed scope to useful client outcomes with consistent quality, the right amount of senior support, and a workload the team can sustain. Assess it across five connected dimensions—delivery flow, quality, client friction, capacity resilience, and consultant development—then read the trends together rather than turning any one measure into a target for an individual.

By Simon ChapmanPublished and reviewed Version 1.1

Download frameworkView measure definitions

Use the numbers to improve how the team works

  • Look at flow and quality together. Faster delivery is not an improvement if the evidence gets worse.
  • Use the median and the spread of results rather than an average that can hide a few badly delayed engagements.
  • Agree when each measure starts and ends, which cases can be excluded, and what counts as serious before comparing periods.
  • Use team-level trends to find problems. Do not turn them into simplistic individual quotas.
  • Review the numbers alongside client feedback and observations from reviewers and delivery staff.

What this framework is—and is not

This is Conversec’s own practical framework. It is not an industry standard or a benchmark. Leaders can adapt its definitions to suit the way their team works. It does not assess an individual tester’s technical skill and should not be used to rank people or teams.

How this relates to the maturity scorecard

The Pentest Team Maturity Scorecard and this framework are two parts of the same improvement conversation, but they do different jobs. The scorecard is a self-assessment of the practices and capabilities a team has in place. This framework uses project, QA, and delivery records to show what is happening over time.

Maturity scorecardPerformance framework
Main questionWhich operating practices and capabilities are in place?What does delivery performance show over time?
EvidenceA structured, honest self-assessment.Project, QA, client, capacity, and development records.
OutputA picture of the team’s strengths, gaps, and next steps.Trends that show where a problem may sit and whether a change is helping.
When to use itTo decide where the team should look more closely.To measure a chosen problem and see whether action is helping.
The different dimension names are intentional. There is no formula for converting a scorecard result into a framework score.

How NIST and CREST inform the framework

NIST SP 800-55 explains how to define, choose, and use measures to understand performance and make improvements. CREST’s penetration testing standards and guidance describe what a well-run service should manage, including preparation, scoping, delivery, reporting, quality assurance, communication, staff competence, and oversight.

SourceHow it informs this framework
NIST SP 800-55 Volume 1 and Volume 2How to choose useful numbers and observations, establish a starting point, check the quality and context of the information, avoid misleading conclusions, and review progress regularly.
CREST Penetration Testing Accreditation Standard and CREST Defensible Penetration Test (2023)What a well-run penetration-testing service needs to manage across scoping, delivery, reporting, quality control, communication, staff competence, limits on the work, oversight, escalation, and sign-off.
Conversec frameworkThe original five dimensions and eight core measures used to understand how the team is performing over time.

These sources help explain what to measure and how to use the results. They do not define Conversec’s dimensions or measures, validate the maturity scorecard, or turn either Conversec tool into an external standard. NIST, CREST, and OWASP have not approved or endorsed this framework.

The five dimensions

Each measure has a primary home: the dimension whose main question it answers most directly. That does not make the relationship exclusive or create a dimension score. Measures can help explain more than one dimension and should still be read together.

DimensionQuestion it answersPrimary measures
Delivery flowDoes work move reliably from scoping through testing and reporting to completion?Scope readiness rate; report turnaround
QualityAre reports technically sound, well evidenced, and ready for review?First-pass acceptance; material QA defect rate
Client frictionWhere do unclear scope, reporting, or explanations create extra client effort?Client clarification volume
Capacity resilienceCan routine work be delivered without repeated senior rescue or hidden extra effort?Senior intervention rate; unbilled delivery load
Consultant developmentAre consultants handling more evidence, judgement, reporting, and client communication independently?Ownership progression

Core measures and definitions

MeasurePrimary dimensionDefinition or formulaRead alongside
Scope readiness rateDelivery flowEngagements meeting agreed prerequisites at planned start ÷ engagements scheduled to startBlocked days and scope-change records
Report turnaroundDelivery flowMedian business days from planned testing end to client-ready draftEngagement complexity and reviewer availability
First-pass acceptanceQualityReports needing no material technical or reasoning change ÷ reports reviewedTypes of serious QA defect
Material QA defect rateQualityReports containing at least one material evidence, accuracy, severity, or recommendation defect ÷ reports reviewedDefect type and recurrence
Client clarification volumeClient frictionDocumented post-draft clarification events per completed engagementCause, stage, and whether clarification was avoidable
Senior intervention rateCapacity resilienceEngagements requiring unplanned senior rescue ÷ completed engagementsReason and time required
Unbilled delivery loadCapacity resilienceUnplanned non-billable recovery hours ÷ total recorded engagement hoursSource of rework or waiting
Ownership progressionConsultant developmentAgreed delivery responsibilities demonstrated independently at the required quality levelExamples of work and reviewer judgement

What counts as a material QA defect?

A material defect is serious enough to change a client decision, make a finding difficult to reproduce, describe risk incorrectly, hide an important limitation, or suggest an unsuitable fix. Track style, optional wording changes, and formatting separately so they do not make quality look worse than it is.

Choosing measures from your scorecard results

Use this table when a scorecard result shows an area worth investigating. One scorecard area may relate to several framework measures. The table helps you choose where to start; it does not convert one score into another.

Scorecard capabilityRelated framework dimensions
Direction and ownershipDelivery flow; capacity resilience
Sales, scoping and handoverDelivery flow; client friction
Delivery disciplineDelivery flow; quality
Technical support and escalationQuality; capacity resilience; consultant development
Evidence, reporting and client communicationQuality; client friction
People, coaching and leadership loadCapacity resilience; consultant development
Improvement and resilienceTrends across all five framework dimensions

Example team dashboard

These are example numbers, not Conversec benchmark data. They show how to read several measures together.

MeasurePrimary dimensionPrevious 12 weeksCurrent 12 weeksInterpretation
Scope readinessDelivery flow72%84%Fewer engagements start without prerequisites.
Median report turnaroundDelivery flow4.5 days3.8 daysFaster reporting is an improvement only if quality has not fallen.
First-pass acceptanceQuality48%61%More work reaches review at the expected standard.
Senior intervention rateCapacity resilience39%24%Routine delivery is becoming less senior-dependent.
Client clarification eventsClient friction1.8 per engagement1.2 per engagementCheck whether the reduction reflects clearer reports.

Start small: a practical way to use the framework

You do not need a new dashboard, perfect data, or all eight measures on day one. Start with one delivery problem that matters to the team and collect only enough evidence to understand it better.

Run a small pilot

  • Name one problem in plain language, such as reports leaving late or senior people repeatedly rescuing routine work.
  • Choose two or three measures that show different sides of that problem.
  • Give each measure a one-sentence definition, an existing data source, and an owner.
  • Check the data for 15 minutes each week. Look for missing or inconsistent records; do not set targets yet.

Build a useful baseline

If the first month produces reliable information, continue long enough to see the team’s normal pattern. Note what else may have affected the results, such as the types of service delivered, unusually complex work, staffing gaps, or changes in record-keeping.

Twelve weeks is a useful starting window, not a rule. Small teams may need longer; high-volume teams may see a stable pattern sooner.

Turn the evidence into one action

  • Review the measures together for about 45 minutes.
  • Ask what changed, what else affected the results, and what is currently causing the most difficulty.
  • Choose one change to try, with an owner and a review date.
  • At the next review, check whether the relevant result improved without another becoming worse.
  • Review definitions quarterly and add a measure only when it helps answer a real decision.

Four terms in plain language

TermWhat it means here
BaselineThe team’s recent, normal pattern before a planned change.
MedianThe middle result after values are ordered. It is less distorted than an average by one exceptionally slow or fast engagement.
ExclusionA case left out for a valid, agreed reason. Record the reason so exclusions do not quietly improve the result.
Material defectA problem serious enough to change a client decision, make a finding difficult to reproduce, describe risk incorrectly, hide a limitation, or suggest the wrong action.

Worked example: late reports and senior rescue

This example shows how to structure a pilot. It is not benchmark data or a claim about a real Conversec client.

DecisionExample
ProblemReports are leaving later than expected, and leaders do not know whether QA rework or senior rescue is the main cause.
MeasuresMedian report turnaround; first-pass acceptance; senior intervention rate.
Existing sourcesProject tracker, QA review record, and a simple note whenever unplanned senior help is required.
Weekly questionAre the records complete enough to show where the delay occurs?
Change to tryUse a quick check of the evidence and report structure before testing ends on complex engagements.
Next reviewAfter four weeks, check whether turnaround and senior intervention improved without material QA defects increasing.

Starter template

Copy these prompts into a shared document or project page. You do not need special measurement software.

PromptWhat to record
Problem we are trying to understandOne specific delivery problem, written in plain language.
MeasuresTwo or three measures that show different sides of the problem.
DefinitionsWhat starts and stops each measure, plus any agreed exclusions.
Sources and ownerWhere the information already lives and who checks it.
What else affected the result?Unusually complex work, staffing changes, missing information, or other facts needed to read the result fairly.
Question to answerThe decision the team wants the information to support.
Next actionOne change to try, its owner, and the date the team will review it.

What the numbers cannot tell you

  • The types of service delivered and the difficulty of engagements can change the results without performance becoming better or worse.
  • Results from a small number of engagements can swing sharply, so show the actual counts as well as percentages.
  • Changes in record-keeping can create false trends.
  • Client questions are not automatically a problem. Separate useful collaboration from questions caused by unclear work.
  • Individual targets can encourage people to improve the number rather than the work. They should not replace professional judgement.

Version history

Version 1.1: Updated external references and evidence base. Replaced NIST SP 800-115 with NIST SP 800-55 Volumes 1 and 2, and added current CREST penetration-testing service guidance.

Version 1.0: Initial publication.

How to cite this framework

Suggested citation: Chapman, Simon. “Pentest Team Performance Framework.” Conversec, version 1.1, 24 August 2026. https://www.conversec.com/frameworks/pentest-team-performance/

Sources and framework status

Conversec created the five-dimension model, measure definitions, and guidance for using them. OWASP reporting guidance is included only as further reading on report quality. No external body has approved or endorsed the framework, and no client or scorecard response data was used.

Start with one delivery problem, not a new reporting programme.

Choose two or three useful measures, learn from the pattern, and make one practical improvement at a time.