Looking for a penetration test? Fixed-scope engagements for GDPR, NIS2 & DORA.Get a quote today →
Request a Quote
Services

AI & LLM Application Penetration Testing

Manual security testing of web applications built on large language models, including chatbots, copilots, RAG search and AI agents, covering the attack classes that traditional web testing does not.

At a Glance
Typical Duration1–2 Weeks
Delivery FormatWritten Report + Walkthrough Call
Best ForChatbots, Copilots, RAG Apps, AI Agents
MethodologyOWASP Top 10 for LLM Applications

Coverage across your AI features

An LLM feature adds new inputs, new trust boundaries and new ways to reach your data. We test the model, and everything it is connected to.

01

Prompt Injection (Direct & Indirect)

Core Coverage

Whether user input, uploaded documents, web pages or emails the model reads can override instructions and change what the application does.

  • Injection paths tested through every input the model consumes, not just the chat box
02

Sensitive Data & System Prompt Leakage

Core Coverage

Whether the model can be coaxed into revealing system prompts, API keys, internal data or other users' information.

  • Responses checked for secrets, internal URLs and personal data across varied phrasing and multi-turn attempts
03

Insecure Output Handling

Core Coverage

What happens when model output is passed on to a browser, database, shell or internal service without validation.

  • Model output traced into downstream sinks, looking for XSS, SSRF, injection and command execution
04

Excessive Agency & Tool Abuse

Core Coverage

Whether the model's tools, plugins and API permissions can be abused to take actions the user was never authorized to perform.

  • Every tool and integration reviewed for least privilege and tested for privilege escalation via the model
05

RAG, Vector Stores & Access Control

Core Coverage

Whether retrieval pipelines respect user and tenant boundaries, or let one user pull another's documents into a response.

  • Cross-user and cross-tenant retrieval tested, plus poisoning of knowledge-base content
06

Guardrails, Abuse & Cost Controls

Core Coverage

How well filters, rate limits and usage caps hold up against jailbreaks, automated abuse and runaway token consumption.

  • Guardrail bypass attempts and resource-exhaustion scenarios against the AI endpoints

What to expect

What's Included

  • Testing of the whole application around the model: authentication, authorization, APIs and session handling
  • Prompt-level attacks run manually against your real system prompts, tools and data sources
  • Review of what the model can reach: connectors, plugins, databases and internal services
  • Testing against staging or a dedicated environment to avoid production side-effects
  • Manual verification of every finding before it reaches the report

Deliverables

  • A fixed price agreed before testing starts
  • Reproduction steps, including the exact prompts and responses, for every finding
  • Severity ratings with practical remediation guidance for prompts, tooling and architecture
  • One retest included once fixes are deployed

Ready to scope an AI test?

Tell us what the AI feature does, which tools and data it can reach, and your timeline, and you'll get a proposal back directly.