Why judgement may become the defining skill in offensive security

As automation expands technical capability, judgement may become the defining pentester skill. Teams need to develop it alongside curiosity and technical depth.

Judgement may become the defining skill in offensive security as automation expands what practitioners can execute. Its value lies in choosing useful objectives, evaluating evidence, setting boundaries and making defensible decisions under uncertainty. Curiosity and technical depth remain essential foundations. Teams also need to preserve the supervised experience through which practitioners learn to exercise that judgement.

By Simon ChapmanPublished Updated Last reviewed

Talk to usAll articles
Offensive security practitioner considering branching system paths, with one route highlighted in gold
Consultant developmentBy Simon Chapman

The short version

  • Greater automated capability increases the importance of deciding what is useful, justified and within scope.
  • Sound judgement depends on technical understanding, curiosity and a willingness to challenge assumptions.
  • Delegation needs clear boundaries, meaningful oversight and identifiable human and organisational accountability.
  • Teams should preserve supervised practice and feedback as automation changes how practitioners gain experience.

Consider a hypothetical penetration test. A tester has demonstrated that one customer can access another customer’s sensitive information. The evidence is clear. An automated workflow offers to continue, collecting more records and exploring related accounts.

The next decision matters. Would continuing answer an outstanding question about the weakness, within the agreed scope? Or would it expose more data without materially improving the client’s understanding?

Stopping might be the right call. So might a carefully bounded follow-up test. The tester needs to understand the evidence, the objective and the consequences well enough to decide.

That is professional judgement. As automation takes on more offensive security work, it may become the skill that most clearly distinguishes an effective practitioner.

What judgement means in offensive security

Judgement is the ability to make a defensible decision from evidence, context and uncertainty while considering the consequences.

It appears throughout an engagement: deciding which attack path deserves attention, whether an unexpected result is meaningful, how much evidence is sufficient and when to involve someone with different expertise.

It also shapes conclusions. A plausible attack path is not necessarily a demonstrated compromise. A technically serious weakness may have consequences that depend on controls the tester has not yet examined. The practitioner needs to recognise those distinctions and explain them.

This is particularly visible when pentest teams handle severity disputes. Good judgement includes changing a conclusion when new evidence warrants it, while resisting pressure to change one without a sound reason.

Experience helps, but seniority does not guarantee sound decisions. Judgement needs evidence, challenge and a willingness to revise assumptions.

Why automation could change the balance

If automation makes more technical execution readily available, the ability to perform an action may become a weaker differentiator between practitioners. Greater weight could fall on choosing the right action and understanding its result.

A tool might generate several promising avenues of investigation. Someone still needs to establish which address the engagement objectives, which rest on unsupported assumptions and which warrant further effort. Pursuing every possibility can consume the time needed to investigate the most consequential one properly.

This is the case for judgement becoming more important. It guides how capability is used.

The change will be uneven. Specialist research and unfamiliar environments can still demand considerable original technical work. There is no need to predict that all execution will become easy, or that every offensive security role will develop in the same way.

But where execution becomes easier to obtain, the quality of the decisions around it may account for more of the value delivered.

Curiosity and technical depth still matter

Curiosity drives investigation. Technical skill helps a practitioner understand systems and test ideas. Judgement determines the direction, the limits and the conclusions.

These capabilities reinforce each other. A tester needs enough technical understanding to recognise when an automated result is implausible, when a test missed a security boundary or when apparently reassuring evidence is incomplete.

Curiosity matters just as much when the initial answer looks convincing. What has been assumed? What alternative explanation fits the observation? What evidence would change the conclusion?

A practitioner who accepts fluent output without asking those questions has little basis for evaluating it. Equally, a technically capable practitioner who pursues an interesting route without considering its relevance can lose sight of the engagement’s purpose.

Calling judgement the defining skill should therefore raise expectations of technical understanding. It takes knowledge to recognise where a tool’s reasoning or evidence falls short.

Deciding what to delegate is part of the job

Automation introduces decisions before testing even begins. Which activities can proceed within agreed boundaries? Which results need verification? Which actions require review before execution?

The answers should reflect the environment and the consequences of error. Organising notes presents different concerns from running a workflow that can change production data. A useful approach is to establish:

  • the objective and authorised boundaries of the delegated work
  • the evidence needed to trust its outputs
  • the conditions that require a pause or escalation
  • who can approve a change and who reviews the outcome

Those boundaries should be practical enough to guide behaviour. Requiring approval without giving the reviewer sufficient context or time provides little assurance.

Nor does judgement always mean being cautious. Sometimes the defensible decision is to investigate further because the current evidence cannot support the conclusion. The point is to understand what the next action would establish and whether it is justified.

Accountability needs people who can explain the decision

There is a strong professional and ethical case for retaining identifiable human and organisational accountability as testing becomes more automated. Clients need people who can explain the decisions made during an engagement and take responsibility for addressing problems.

The NIST AI Risk Management Framework’s governance guidance calls for defined responsibilities, executive responsibility for AI risk decisions and policies covering human oversight. This is voluntary risk-management guidance, rather than a universal legal requirement for a human to approve every action.

CREST’s principles for AI-enabled activities, updated in May 2026, place oversight with suitably competent personnel who can review outputs, challenge decisions and intervene when needed. They also call for AI activity to remain within defined controls and, where relevant, client-agreed scope. For offensive security teams, this makes the ability to evaluate and govern delegated work a practical professional responsibility.

The practical question is whether the people responsible understand and can govern what has been delegated. Legal obligations depend on the applicable jurisdiction and circumstances; a broad claim that the law will always reserve judgement to humans would go too far.

Professional responsibility also involves explaining the reasoning to the client. This is one reason technical pentesters need consulting skills: a decision becomes useful when its evidence, limitations and implications can be understood.

How will practitioners develop the judgement they need?

There is a potential tension here. Teams may need stronger judgement while automating some of the work through which people previously developed it.

Manually investigating an inconsistent result can teach a tester how a system behaves. Writing a finding forces them to connect evidence to a claim. Discussing a mistake with a reviewer can expose an assumption they did not know they were making.

If tools take over those activities, the output may improve while opportunities to practise the underlying reasoning shrink. That is a development risk to manage, not an inevitable outcome. Automation can also support learning by making it easier to explore alternatives and obtain feedback.

This builds on how AI changes the junior pentester role, but the issue extends to experienced practitioners working in unfamiliar environments.

Teams can preserve useful learning deliberately:

  • Ask testers to propose a next step and explain their reasoning before comparing it with a tool’s suggestion.
  • Keep supervised hands-on testing in development plans, including uncertain and misleading results.
  • Review decisions after engagements, including good decisions to stop, escalate or change direction.
  • Use short scenarios where evidence, client pressure and operational constraints pull in different directions.

The aim is to preserve the experiences that develop understanding. Repeating every mechanical task manually is unnecessary; ensuring practitioners can evaluate the work remains essential.

This is why technical training does not fix every pentest team problem. People also need coached practice and feedback on the decisions they make. Good pentest QA can provide that feedback when reviewers examine reasoning and discuss it with the tester.

Make judgement something teams develop and assess

For practitioners, a useful habit is to make significant decisions explainable: what was known, what remained uncertain, why the action was proportionate and what would have changed the decision.

For leaders, development and assessment should make room for that reasoning alongside technical results. Can someone challenge an automated recommendation? Recognise insufficient evidence? Adapt to new information? Explain a decision under client scrutiny?

Judgement may emerge as the most important skill in offensive security. The argument does not require certainty about that ranking. As automation expands capability, teams have good reason to invest in the people who decide how that capability should be used.

Sources and further reading

External standards and guidance support the factual context in this article. The analysis and recommendations are Conversec practitioner guidance unless attributed otherwise.

If this article describes a real delivery pressure, turn it into a next step.

Conversec helps offensive security teams improve consulting maturity, leadership capacity, and delivery clarity.