PrivSec Consulting
  • Home
  • About
  • Services
    • Governance, Risk & Compliance
    • Penetration Testing
    • Configuration Reviews
    • Code Review
    • Privacy
    • Artificial Intelligence
    • Security Resilience Improvement Exercises
    • Security Awareness and Training
    • Alignment and Uplift Activities >
      • PCI DSS
    • Consultancy and Advice
  • Releases
  • Contact

Releases

Give an app to 5 different security testers, you'll get 5 different reports

5/19/2026

 
If you've commissioned more than one penetration test over the years, you've probably noticed that different testers, and different vendors, will all give you different results, even on the same app. Maybe the second test found things the first missed. Maybe the two reports barely overlap. You might have even wondered if the testers were looking at the same application.

You're not imagining it, and it's not always a sign something went wrong.

At NZ Tech Rally this week I walked through the main reasons this happens, and what you can do to get better outcomes from your testing programme. Here's the written version.

The people factor
Penetration testing involves human judgement, and human judgement varies from person to person. That's the starting point for everything else.

A web application specialist will have a different focus from the generalist vs the tester who has a focus on network pentesting. A tester who found authentication bypasses last week will look at auth harder this week. One who trained through bug bounty programmes has developed pattern recognition very different from one who came through a university computer science programme or development background.

The assumed threat model matters too. Some testers approach an engagement asking "what would an insider do?" Others are thinking "what does the script kiddy do?" Some follow a checklist methodically, and others start adversarially and see where the thread leads. Neither approach is wrong, they just find different things.

Risk appetite adds another layer. One tester marks a finding medium severity, and another calls it high. Severity is partly subjective, shaped by context, sector experience, and professional background.

None of this is unique to security. Two financial auditors reviewing the same accounts will flag different items as material risks depending on their background and professional scepticism. Five UX researchers evaluating the same interface produce five different pain point lists. Experienced clinicians reviewing identical patient data reach different diagnoses at measurable rates, even specialists.

Methodology and tooling
How a tester approaches the problem shapes what they find. Different methodologies will lead to different findings. Some will ensure better coverage, some will find more creative issues.

The AI shift is changing this landscape in real time. AI assisted tools are generating findings faster than ever. They're also surfacing false positives that require triage, and they can miss logic flaws entirely. Automated tools find known patterns. Manual testing finds unknown logic flaws. A good engagement uses both, and your tester should know the best place for both.

Scope is doing more work than you think
This is the first place I look when clients are surprised by differing results between engagements.

"Test our website" and "Test our application at [URL], authenticated and unauthenticated, including APIs, mobile clients, and admin surfaces, with business logic and trust boundaries defined" are not the same brief. They produce completely different engagements. Vague scope doesn't just create ambiguity, it creates implicit decisions about what gets tested, made by the tester rather than by you.
A good scope may look like:
  • A named application, URL, and environments
  • Explicit coverage of authenticated and unauthenticated access (or whatever is suitable in your environment, for your app)
  • APIs, mobile, and admin surfaces included
  • Business logic and trust boundaries defined upfront
  • Agreement on white box versus black box approach
  • A shared understanding of what a successful attack would actually look like
If you've never had a conversation with your tester about what "success" means for an attacker in your context, that's a gap worth closing before the next engagement starts.

The reality of constraints
A short engagement forces prioritisation, that a larger scope may not. Whatever gets tested first gets tested most thoroughly. Features scoped late in the process get much less effort given to them.

"We found nothing critical" can mean exactly that, or it can mean "we didn't have time to look deeper." You need to be considerate of that when providing guidance around how much effort you'd like given to the testing.

How can you address that? Tell your tester explicitly what to prioritise if time runs short. "If you run out of time, focus on authentication, payment flows, and data exports." Don't let scope prioritisation be at the mercy of the tester.

Security is a moving target
A clean report today doesn't mean a clean report tomorrow. Log4Shell (CVE-2021-44228, CVSS 10.0) arrived in December 2021. A clean pentest in November meant nothing by December, as the library was in use across thousands of applications. The XZ Utils backdoor in March 2024 was a supply chain attack buried in open source that no pentest would have caught. It needed code review, and was ultimately found by someone benchmarking performance who noticed something unusual. React2Shell (CVE-2025-55182, CVSS 10.0) dropped in December 2025 with exploitation observed within hours of disclosure.

None of these issues were highlighted in services that leveraged them... until they were. Penetration testing is a point-in-time snapshot. It tells you about the security posture of your application at the moment the test was conducted, against the knowledge and techniques available at that time. It does not tell you about vulnerabilities discovered the following month, or about changes introduced during your next deployment cycle. Investing regularly in testing is the best way to give yourselves ongoing confidence in your security posture, especially for external assets. Treat it as one layer of your security programme, not the whole picture.

What you can actually do?
Define scope collaboratively. Share architecture diagrams, data flow documentation, and business context before the engagement starts. The tester who understands your application is a more effective tester.

Be explicit about priorities. If time runs short, what matters most? Authentication, payment flows, data exports?

Treat reports as conversations. Debrief after the engagement. Ask what they would have looked at with more time. Ask what they found interesting that didn't make the final report. Sometimes there's interesting paths a tester went down that never made it to the report, as it didn't end in exploitation, however that may have just been a limitation of the test due to time.

Retest, don't just remediate. Fixes introduce bugs. Verify remediations with a targeted retest rather than just internal validation.

Layer your controls.
Penetration testing alongside a bug bounty programme, SAST in your pipeline, solid architecture, or even a WAF will give you overlapping coverage. No single control catches everything, and layered controls provide much more confidence.

This summarises a talk our Managing Director, Peter Jakowetz did at the NZ Tech Rally in May 2026.



Comments are closed.

Want to know more? Contact us now.

[email protected] | 0800 150 805
  • Home
  • About
  • Services
    • Governance, Risk & Compliance
    • Penetration Testing
    • Configuration Reviews
    • Code Review
    • Privacy
    • Artificial Intelligence
    • Security Resilience Improvement Exercises
    • Security Awareness and Training
    • Alignment and Uplift Activities >
      • PCI DSS
    • Consultancy and Advice
  • Releases
  • Contact