How to use the framework
Use the numbers to improve how the team works
- Look at flow and quality together. Faster delivery is not an improvement if the evidence gets worse.
- Use the median and the spread of results rather than an average that can hide a few badly delayed engagements.
- Agree when each measure starts and ends, which cases can be excluded, and what counts as serious before comparing periods.
- Use team-level trends to find problems. Do not turn them into simplistic individual quotas.
- Review the numbers alongside client feedback and observations from reviewers and delivery staff.
What this framework is—and is not
This is Conversec’s own practical framework. It is not an industry standard or a benchmark. Leaders can adapt its definitions to suit the way their team works. It does not assess an individual tester’s technical skill and should not be used to rank people or teams.
How this relates to the maturity scorecard
The Pentest Team Maturity Scorecard and this framework are two parts of the same improvement conversation, but they do different jobs. The scorecard is a self-assessment of the practices and capabilities a team has in place. This framework uses project, QA, and delivery records to show what is happening over time.
| Maturity scorecard | Performance framework | |
|---|---|---|
| Main question | Which operating practices and capabilities are in place? | What does delivery performance show over time? |
| Evidence | A structured, honest self-assessment. | Project, QA, client, capacity, and development records. |
| Output | A picture of the team’s strengths, gaps, and next steps. | Trends that show where a problem may sit and whether a change is helping. |
| When to use it | To decide where the team should look more closely. | To measure a chosen problem and see whether action is helping. |
The different dimension names are intentional. There is no formula for converting a scorecard result into a framework score.
How NIST and CREST inform the framework
NIST SP 800-55 explains how to define, choose, and use measures to understand performance and make improvements. CREST’s penetration testing standards and guidance describe what a well-run service should manage, including preparation, scoping, delivery, reporting, quality assurance, communication, staff competence, and oversight.
| Source | How it informs this framework |
|---|---|
| NIST SP 800-55 Volume 1 and Volume 2 | How to choose useful numbers and observations, establish a starting point, check the quality and context of the information, avoid misleading conclusions, and review progress regularly. |
| CREST Penetration Testing Accreditation Standard and CREST Defensible Penetration Test (2023) | What a well-run penetration-testing service needs to manage across scoping, delivery, reporting, quality control, communication, staff competence, limits on the work, oversight, escalation, and sign-off. |
| Conversec framework | The original five dimensions and eight core measures used to understand how the team is performing over time. |
These sources help explain what to measure and how to use the results. They do not define Conversec’s dimensions or measures, validate the maturity scorecard, or turn either Conversec tool into an external standard. NIST, CREST, and OWASP have not approved or endorsed this framework.
The five dimensions
Each measure has a primary home: the dimension whose main question it answers most directly. That does not make the relationship exclusive or create a dimension score. Measures can help explain more than one dimension and should still be read together.
| Dimension | Question it answers | Primary measures |
|---|---|---|
| Delivery flow | Does work move reliably from scoping through testing and reporting to completion? | Scope readiness rate; report turnaround |
| Quality | Are reports technically sound, well evidenced, and ready for review? | First-pass acceptance; material QA defect rate |
| Client friction | Where do unclear scope, reporting, or explanations create extra client effort? | Client clarification volume |
| Capacity resilience | Can routine work be delivered without repeated senior rescue or hidden extra effort? | Senior intervention rate; unbilled delivery load |
| Consultant development | Are consultants handling more evidence, judgement, reporting, and client communication independently? | Ownership progression |
Core measures and definitions
| Measure | Primary dimension | Definition or formula | Read alongside |
|---|---|---|---|
| Scope readiness rate | Delivery flow | Engagements meeting agreed prerequisites at planned start ÷ engagements scheduled to start | Blocked days and scope-change records |
| Report turnaround | Delivery flow | Median business days from planned testing end to client-ready draft | Engagement complexity and reviewer availability |
| First-pass acceptance | Quality | Reports needing no material technical or reasoning change ÷ reports reviewed | Types of serious QA defect |
| Material QA defect rate | Quality | Reports containing at least one material evidence, accuracy, severity, or recommendation defect ÷ reports reviewed | Defect type and recurrence |
| Client clarification volume | Client friction | Documented post-draft clarification events per completed engagement | Cause, stage, and whether clarification was avoidable |
| Senior intervention rate | Capacity resilience | Engagements requiring unplanned senior rescue ÷ completed engagements | Reason and time required |
| Unbilled delivery load | Capacity resilience | Unplanned non-billable recovery hours ÷ total recorded engagement hours | Source of rework or waiting |
| Ownership progression | Consultant development | Agreed delivery responsibilities demonstrated independently at the required quality level | Examples of work and reviewer judgement |
What counts as a material QA defect?
A material defect is serious enough to change a client decision, make a finding difficult to reproduce, describe risk incorrectly, hide an important limitation, or suggest an unsuitable fix. Track style, optional wording changes, and formatting separately so they do not make quality look worse than it is.
Choosing measures from your scorecard results
Use this table when a scorecard result shows an area worth investigating. One scorecard area may relate to several framework measures. The table helps you choose where to start; it does not convert one score into another.
| Scorecard capability | Related framework dimensions |
|---|---|
| Direction and ownership | Delivery flow; capacity resilience |
| Sales, scoping and handover | Delivery flow; client friction |
| Delivery discipline | Delivery flow; quality |
| Technical support and escalation | Quality; capacity resilience; consultant development |
| Evidence, reporting and client communication | Quality; client friction |
| People, coaching and leadership load | Capacity resilience; consultant development |
| Improvement and resilience | Trends across all five framework dimensions |
Example team dashboard
These are example numbers, not Conversec benchmark data. They show how to read several measures together.
| Measure | Primary dimension | Previous 12 weeks | Current 12 weeks | Interpretation |
|---|---|---|---|---|
| Scope readiness | Delivery flow | 72% | 84% | Fewer engagements start without prerequisites. |
| Median report turnaround | Delivery flow | 4.5 days | 3.8 days | Faster reporting is an improvement only if quality has not fallen. |
| First-pass acceptance | Quality | 48% | 61% | More work reaches review at the expected standard. |
| Senior intervention rate | Capacity resilience | 39% | 24% | Routine delivery is becoming less senior-dependent. |
| Client clarification events | Client friction | 1.8 per engagement | 1.2 per engagement | Check whether the reduction reflects clearer reports. |
Start small: a practical way to use the framework
You do not need a new dashboard, perfect data, or all eight measures on day one. Start with one delivery problem that matters to the team and collect only enough evidence to understand it better.
Phase one · first 30 days
Run a small pilot
- Name one problem in plain language, such as reports leaving late or senior people repeatedly rescuing routine work.
- Choose two or three measures that show different sides of that problem.
- Give each measure a one-sentence definition, an existing data source, and an owner.
- Check the data for 15 minutes each week. Look for missing or inconsistent records; do not set targets yet.
Phase two · up to 12 weeks
Build a useful baseline
If the first month produces reliable information, continue long enough to see the team’s normal pattern. Note what else may have affected the results, such as the types of service delivered, unusually complex work, staffing gaps, or changes in record-keeping.
Twelve weeks is a useful starting window, not a rule. Small teams may need longer; high-volume teams may see a stable pattern sooner.
Phase three · monthly
Turn the evidence into one action
- Review the measures together for about 45 minutes.
- Ask what changed, what else affected the results, and what is currently causing the most difficulty.
- Choose one change to try, with an owner and a review date.
- At the next review, check whether the relevant result improved without another becoming worse.
- Review definitions quarterly and add a measure only when it helps answer a real decision.
Four terms in plain language
| Term | What it means here |
|---|---|
| Baseline | The team’s recent, normal pattern before a planned change. |
| Median | The middle result after values are ordered. It is less distorted than an average by one exceptionally slow or fast engagement. |
| Exclusion | A case left out for a valid, agreed reason. Record the reason so exclusions do not quietly improve the result. |
| Material defect | A problem serious enough to change a client decision, make a finding difficult to reproduce, describe risk incorrectly, hide a limitation, or suggest the wrong action. |
Worked example: late reports and senior rescue
This example shows how to structure a pilot. It is not benchmark data or a claim about a real Conversec client.
| Decision | Example |
|---|---|
| Problem | Reports are leaving later than expected, and leaders do not know whether QA rework or senior rescue is the main cause. |
| Measures | Median report turnaround; first-pass acceptance; senior intervention rate. |
| Existing sources | Project tracker, QA review record, and a simple note whenever unplanned senior help is required. |
| Weekly question | Are the records complete enough to show where the delay occurs? |
| Change to try | Use a quick check of the evidence and report structure before testing ends on complex engagements. |
| Next review | After four weeks, check whether turnaround and senior intervention improved without material QA defects increasing. |
Starter template
Copy these prompts into a shared document or project page. You do not need special measurement software.
| Prompt | What to record |
|---|---|
| Problem we are trying to understand | One specific delivery problem, written in plain language. |
| Measures | Two or three measures that show different sides of the problem. |
| Definitions | What starts and stops each measure, plus any agreed exclusions. |
| Sources and owner | Where the information already lives and who checks it. |
| What else affected the result? | Unusually complex work, staffing changes, missing information, or other facts needed to read the result fairly. |
| Question to answer | The decision the team wants the information to support. |
| Next action | One change to try, its owner, and the date the team will review it. |
What the numbers cannot tell you
- The types of service delivered and the difficulty of engagements can change the results without performance becoming better or worse.
- Results from a small number of engagements can swing sharply, so show the actual counts as well as percentages.
- Changes in record-keeping can create false trends.
- Client questions are not automatically a problem. Separate useful collaboration from questions caused by unclear work.
- Individual targets can encourage people to improve the number rather than the work. They should not replace professional judgement.
Version history
Version 1.1: Updated external references and evidence base. Replaced NIST SP 800-115 with NIST SP 800-55 Volumes 1 and 2, and added current CREST penetration-testing service guidance.
Version 1.0: Initial publication.
How to cite this framework
Suggested citation: Chapman, Simon. “Pentest Team Performance Framework.” Conversec, version 1.1, 24 August 2026. https://www.conversec.com/frameworks/pentest-team-performance/
Evidence
Sources and framework status
- NIST SP 800-55 Volume 1: Identifying and Selecting Measures
- NIST SP 800-55 Volume 2: Developing an Information Security Measurement Program
- CREST Accreditation Standards: current Penetration Testing Accreditation Standard
- CREST Defensible Penetration Test (2023)
- OWASP Web Security Testing Guide: Reporting Structure (supplementary reporting reference)
Conversec created the five-dimension model, measure definitions, and guidance for using them. OWASP reporting guidance is included only as further reading on report quality. No external body has approved or endorsed the framework, and no client or scorecard response data was used.
