Skip to content
Vaptiq logo mark — V orientationVAPTIQ
All services

[ Adversarial ]

AI & LLM Penetration Testing

Prompt injection, tool abuse and data leakage in systems that act on their own output.
Typical duration
5-10 days
Starting price
Scoped
You provide
System access, tool definitions, prompt architecture

[ 01 / The engagement ]

An LLM with tool access is a new kind of confused deputy: it reads untrusted text and then takes privileged actions based on it. We test the whole system rather than the model, the retrieval corpus, the tool definitions, the agent loop and the permissions behind it, because the damaging failures happen when injected instructions reach something that can write, spend or send.

[ 02 / Coverage ]

What we test, in practice.

This is the working checklist, not a marketing list. Anything your scope adds gets written into the engagement letter before we start.

Methodology

OWASP Top 10 for LLM ApplicationsMITRE ATLASNIST AI RMF
  1. 01Direct and indirect prompt injection through retrieved content
  2. 02Tool and function-call abuse, including chained privileged actions
  3. 03Training and RAG corpus poisoning
  4. 04System prompt extraction and guardrail bypass
  5. 05Sensitive data leakage through model and vector-store responses
  6. 06Agent loop manipulation and unbounded autonomous action
  7. 07Excessive agency review of the permissions behind each tool

[ 03 / What you get ]

Four things land at the end of every engagement.

Technical report

Every finding with CVSS v4.0 score, evidence, reproduction steps and a specific fix, written for the engineer who has to close it.

Executive summary

Two pages your board can read. Risk in business terms, with the three things that matter most called out.

Letter of attestation

A shareable document proving the test happened and what it covered, for customers and auditors who should not see the full report.

Free retest

Once you have fixed things, the same tester verifies each finding and reissues the report. Included for 90 days.

[ 04 / Questions ]

Before you commit.

Only where a jailbreak has consequences. Making a model say something rude is not a finding. Making it call an internal tool with someone else's authorisation is, and that is what we test for.

Less than you would think. Almost every serious finding sits in the surrounding architecture, what the agent is allowed to do, and what untrusted text can reach it, not in the model weights.

Scope a ai & llm penetration testing.

Send us the target and the deadline. You get a written scope and a fixed price, usually within one working day.