# Pentest Team Performance Framework

Version 1.1 — published and reviewed 24 August 2026

Author: Simon Chapman, Conversec

Canonical HTML version: https://www.conversec.com/frameworks/pentest-team-performance/

## Definition

Pentest team performance means moving work from a clear, agreed scope to useful client outcomes with consistent quality, the right amount of senior support, and a workload the team can sustain.

## How this relates to the Pentest Team Maturity Scorecard

The scorecard and framework are complementary, but they answer different questions:

| Tool | Main question | Evidence | Output |
| --- | --- | --- | --- |
| Pentest Team Maturity Scorecard | Which operating practices and capabilities are in place? | A structured self-assessment across seven capability areas | A picture of the team's strengths, gaps, and next steps |
| Pentest Team Performance Framework | What does delivery performance show over time? | Project, QA, client, capacity, and development records across five dimensions | Trends that show where a problem may sit and whether a change is helping |

The different dimension names are intentional. A scorecard result does not convert into a framework score. Use the scorecard to choose where to look and the framework to choose what to measure.

## How NIST and CREST inform the framework

NIST SP 800-55 explains how to define, choose, and use measures to understand performance and make improvements. CREST's penetration testing standards and guidance describe what a well-run service should manage, including preparation, scoping, delivery, reporting, quality assurance, communication, staff competence, and oversight.

| Source | How it informs this framework |
| --- | --- |
| [NIST SP 800-55 Volume 1](https://csrc.nist.gov/pubs/sp/800/55/v1/final) and [Volume 2](https://csrc.nist.gov/pubs/sp/800/55/v2/final) | How to choose useful numbers and observations, establish a starting point, check the quality and context of the information, avoid misleading conclusions, and review progress regularly |
| [CREST Penetration Testing Accreditation Standard](https://www.crest-approved.org/membership/accreditation-standards/) and [CREST Defensible Penetration Test (2023)](https://www.crest-approved.org/wp-content/uploads/2022/12/CREST-Defensible-Penetration-Test-v5-2.pdf) | What a well-run penetration-testing service needs to manage across scoping, delivery, reporting, quality control, communication, staff competence, limits on the work, oversight, escalation, and sign-off |
| Conversec framework | The original five dimensions and eight core measures used to understand how the team is performing over time |

These sources help explain what to measure and how to use the results. They do not define Conversec's dimensions or measures, validate the maturity scorecard, or turn either Conversec tool into an external standard. NIST, CREST, and OWASP have not approved or endorsed this framework.

## Five dimensions

Each measure has a primary home: the dimension whose main question it answers most directly. That does not make the relationship exclusive or create a dimension score. Measures can help explain more than one dimension and should still be read together.

| Dimension | Question it answers | Primary measures |
| --- | --- | --- |
| Delivery flow | Does work move reliably from scoping through testing and reporting to completion? | Scope readiness rate; report turnaround |
| Quality | Are reports technically sound, well evidenced, and ready for review? | First-pass acceptance; material QA defect rate |
| Client friction | Where do unclear scope, reporting, or explanations create extra client effort? | Client clarification volume |
| Capacity resilience | Can routine work be delivered without repeated senior rescue or hidden extra effort? | Senior intervention rate; unbilled delivery load |
| Consultant development | Are consultants handling more evidence, judgement, reporting, and client communication independently? | Ownership progression |

## Core measures and definitions

| Measure | Primary dimension | Definition or formula | Read alongside |
| --- | --- | --- | --- |
| Scope readiness rate | Delivery flow | Engagements meeting agreed prerequisites at planned start / engagements scheduled to start | Blocked days and scope-change records |
| Report turnaround | Delivery flow | Median business days from planned testing end to client-ready draft | Engagement complexity and reviewer availability |
| First-pass acceptance | Quality | Reports needing no material technical or reasoning change / reports reviewed | Types of serious QA defect |
| Material QA defect rate | Quality | Reports with at least one material evidence, accuracy, severity, or recommendation defect / reports reviewed | Defect type and recurrence |
| Client clarification volume | Client friction | Documented post-draft clarification events per completed engagement | Cause, stage, and whether clarification was avoidable |
| Senior intervention rate | Capacity resilience | Engagements requiring unplanned senior rescue / completed engagements | Reason and time required |
| Unbilled delivery load | Capacity resilience | Unplanned non-billable recovery hours / total recorded engagement hours | Source of rework or waiting |
| Ownership progression | Consultant development | Agreed responsibilities demonstrated independently at the required quality level | Examples of work and reviewer judgement |

## Choosing measures from your scorecard results

One scorecard area may relate to several framework measures. This table helps you choose where to start; it does not convert one score into another.

| Scorecard capability | Related framework dimensions |
| --- | --- |
| Direction and ownership | Delivery flow; capacity resilience |
| Sales, scoping and handover | Delivery flow; client friction |
| Delivery discipline | Delivery flow; quality |
| Technical support and escalation | Quality; capacity resilience; consultant development |
| Evidence, reporting and client communication | Quality; client friction |
| People, coaching and leadership load | Capacity resilience; consultant development |
| Improvement and resilience | Trends across all five framework dimensions |

## Start small: a practical implementation guide

You do not need a new dashboard, perfect data, or all eight measures on day one. Start with one delivery problem and enough evidence to understand it better.

### Phase one: first 30 days

1. Name one problem in plain language.
2. Choose two or three measures that show different sides of it.
3. Give each measure a one-sentence definition, an existing data source, and an owner.
4. Check the data for 15 minutes each week. Look for missing or inconsistent records; do not set targets yet.

### Phase two: build a baseline

If the first month produces reliable information, continue long enough to see the normal pattern. Note what else may have affected the results, such as the types of service delivered, unusually complex work, staffing gaps, or changes in record-keeping. Twelve weeks is a useful starting window, not a rule.

### Phase three: monthly improvement rhythm

1. Review the measures together for about 45 minutes.
2. Ask what changed, what else affected the results, and what is currently causing the most difficulty.
3. Choose one change to try, with an owner and review date.
4. At the next review, check whether the relevant result improved without another becoming worse.
5. Review definitions quarterly and add measures only when they answer a real decision.

### Plain-language terms

| Term | Meaning |
| --- | --- |
| Baseline | The team's recent, normal pattern before a planned change |
| Median | The middle result after values are ordered; it is less distorted than an average by one unusual engagement |
| Exclusion | A case left out for a valid, agreed, and recorded reason |
| Material defect | A problem serious enough to change a client decision, make a finding difficult to reproduce, describe risk incorrectly, hide a limitation, or suggest the wrong action |

### Worked example: late reports and senior rescue

| Decision | Example |
| --- | --- |
| Problem | Reports are leaving later than expected, and leaders do not know whether QA rework or senior rescue is the main cause |
| Measures | Median report turnaround; first-pass acceptance; senior intervention rate |
| Sources | Project tracker, QA review record, and a note whenever unplanned senior help is required |
| Weekly question | Are the records complete enough to show where the delay occurs? |
| Change to try | Use a quick check of the evidence and report structure before testing ends on complex engagements |
| Next review | After four weeks, check whether turnaround and senior intervention improved without material QA defects increasing |

This example shows how to structure a pilot. It is not benchmark data or a claim about a real Conversec client.

### Starter template

| Prompt | What to record |
| --- | --- |
| Problem we are trying to understand | One specific delivery problem, written in plain language |
| Measures | Two or three measures that show different sides of the problem |
| Definitions | What starts and stops each measure, plus agreed exclusions |
| Sources and owner | Where the information already lives and who checks it |
| What else affected the result? | Unusually complex work, staffing changes, missing information, or other facts needed to read the result fairly |
| Question to answer | The decision the team wants the information to support |
| Next action | One change to try, its owner, and the date the team will review it |

## What the numbers cannot tell you

- The types of service delivered and the difficulty of engagements can change the results without performance becoming better or worse.
- Results from a small number of engagements can swing sharply, so show the actual counts as well as percentages.
- Changes in record-keeping can create false trends.
- Client questions are not automatically a problem. Separate useful collaboration from questions caused by unclear work.
- Individual targets can encourage people to improve the number rather than the work. They should not replace professional judgement.

This is Conversec's own practical guidance, not an industry standard or benchmark. No client or scorecard response data was used.

## Sources and further reading

- [NIST SP 800-55 Volume 1: Identifying and Selecting Measures](https://csrc.nist.gov/pubs/sp/800/55/v1/final)
- [NIST SP 800-55 Volume 2: Developing an Information Security Measurement Program](https://csrc.nist.gov/pubs/sp/800/55/v2/final)
- [CREST Accreditation Standards: current Penetration Testing Accreditation Standard](https://www.crest-approved.org/membership/accreditation-standards/)
- [CREST Defensible Penetration Test (2023)](https://www.crest-approved.org/wp-content/uploads/2022/12/CREST-Defensible-Penetration-Test-v5-2.pdf)
- [OWASP Web Security Testing Guide: Reporting Structure](https://wstg.owasp.org/latest/5-Reporting/01-Reporting_Structure/) — supplementary reporting reference only

Conversec created the five-dimension model, measure definitions, and guidance for using them. No external body has approved or endorsed the framework.

## Version history

- **Version 1.1:** Updated external references and evidence base. Replaced NIST SP 800-115 with NIST SP 800-55 Volumes 1 and 2, and added current CREST penetration-testing service guidance.
- **Version 1.0:** Initial publication.

Suggested citation: Chapman, Simon. "Pentest Team Performance Framework." Conversec, version 1.1, 24 August 2026. https://www.conversec.com/frameworks/pentest-team-performance/
