Skip to content

Services, Wiki-Artikel und Blog-Beiträge durchsuchen

↑↓NavigierenEnterÖffnenESCSchließen

AI Penetration Testing

How secure is your Artificial Intelligence?

Prompt injection. Jailbreaking. Data exfiltration. We test your LLMs, RAG systems, and AI agents the way real attackers do - per OWASP Top 10 LLM and MITRE ATLAS.

OWASP Top 10 LLM MITRE ATLAS EU AI Act Art. 15 ISO 42001

Trusted by our clients

OWASP LLM Top 10 Categories
10
Pentests completed
500+
Fixed-price quote (business days)
48h
Subcontractors
0

The Problem

AI systems are being attacked - differently from traditional software

Your LLM chatbot, your AI copilot, your automated decision logic - they all have an attack surface that no classical penetration test covers. Prompt injection alone affects every LLM application. And the regulatory clock is ticking:

EU AI Act - Article 15

High-risk AI must demonstrably be robust against adversarial attacks. GPAI governance applies since August 2025.

NIS-2 & DORA

AI-powered systems in critical infrastructure and financial services are subject to the same security audit requirements - with personal liability for management.

GDPR Risk

LLMs can expose trained personal data. A single data leakage incident can trigger regulatory fines and reputational damage.

TRADITIONAL PENTEST

Tests networks, APIs, web apps, infrastructure - but not AI logic, model behavior, or guardrails.

OWASP Top 10 PTES ATT&CK

AI PENETRATION TEST

Additionally tests: prompt injection, jailbreaking, data exfiltration, guardrail bypass, agent behavior, model integrity, RAG poisoning.

OWASP Top 10 LLM MITRE ATLAS NIST AI RMF EU AI Act ISO 42001

Methodology

Our five-phase process

01

2-3 days

Scoping & Threat Modeling

Identification of all AI components, threat modeling per MITRE ATLAS, definition of rules of engagement and test scope.

02

3-5 days

Reconnaissance

Analysis of AI architecture: model endpoints, API interfaces, data pipelines, guardrail configuration, agent capabilities, and integrations.

03

5-10 days

Vulnerability Testing

Automated scans (Garak, Promptfoo) combined with manual expert analysis. Systematic testing of all OWASP Top 10 LLM categories and MITRE ATLAS techniques.

04

2-5 days

Exploitation & PoC

Confirmation of critical findings with proof-of-concept. Chaining vulnerabilities into realistic attack scenarios with quantified business impact.

05

2-4 days

Reporting & Remediation

Technical report with CVSS scoring, compliance mapping (OWASP, EU AI Act, ISO 42001, NIST AI RMF), and prioritized remediation roadmap. Management summary and closing presentation.

Typical total duration: 15-25 days - depending on scope and number of AI components.
You receive a binding fixed-price quote within 48 business hours.

Compliance

One test - all evidence

Every finding is mapped to the relevant standards. Your report is audit-ready.

OWASP Top 10 LLM

Systematic testing of all 10 vulnerability categories for LLM applications - the de facto standard for LLM security.

International community · Open source

MITRE ATLAS

Threat modeling and attack scenarios per the AI-specific counterpart to MITRE ATT&CK.

Tactics · Techniques · Procedures

EU AI Act

Evidence of requirements from Article 15: accuracy, robustness, cybersecurity for high-risk AI.

Art. 15 · GPAI since Aug. 2025

ISO/IEC 42001

Technical evidence for the controls of the AI management system standard - basis for certification.

38 controls · 9 objectives

NIST AI RMF

Mapping to the four core functions Govern, Map, Measure, Manage of the AI Risk Management Framework.

Incl. GenAI Profile (2024)

NIS-2 / DORA

Integration into existing NIS-2 security requirements and DORA threat-led penetration testing obligations for financial entities.

Critical infrastructure · Financial sector

Packages

Transparent pricing

Fixed-price quotes within 48 business hours. No hourly rates, no surprises.

FOCUSED

LLM Pentest

Single chatbot or copilot

from EUR 8,100excl. VAT

  • Full OWASP Top 10 LLM
  • Prompt injection & jailbreaking
  • Data exfiltration tests
  • Guardrail bypass assessment
  • Technical report + management summary
Get a quote
Recommended

COMPREHENSIVE

AI Security Assessment

Multiple AI components + RAG

from EUR 14,850excl. VAT

  • Everything in LLM Pentest
  • RAG system security
  • AI agent testing
  • ML model review
  • Compliance mapping (EU AI Act, ISO 42001)
  • Closing presentation + workshop
Get a quote

PREMIUM

AI Red Teaming

Adversarial simulation · 4-6 weeks

from EUR 25,650excl. VAT

  • Everything in AI Security Assessment
  • Creative attack scenarios
  • Multi-vector exploitation
  • Realistic threat simulation
  • Continuous testing over weeks
  • Purple team debrief
Get a quote

Why AWARE7

What sets us apart from other providers

Pure awareness platforms don't test systems. Pure consulting firms are too far removed. AWARE7 combines both: we hack your infrastructure and train your employees: tailored to mid-sized companies, personal, without enterprise overhead.

Research and teaching as our foundation

20 %

Around 20% of our revenue comes from research projects for the BSI and the BMBF. Our studies, published at ACM and Springer conferences, analyse millions of websites and tens of thousands of phishing emails. Three of our executives are professors at German universities at the same time.

Digital sovereignty: no compromises

100 %

All data is stored and processed exclusively in Germany, without US cloud providers. All staff are permanently employed, covered by social insurance and bound by uniform legal obligations.

Fixed price within 24h: predictable project timelines

24 h

Within 24 hours you receive a binding fixed-price quote without hourly rate risk. A well-practised team and standardised processes ensure a clear schedule with a defined start and end date.

Your dedicated contact

1:1

A personal project manager accompanies you from the first meeting to the retest. You book appointments directly with your contact person and keep the same contact throughout the project.

Peer-reviewed publications

First page of the paper: Different Seas, Different Phishes - Large-Scale Analysis of Phishing Simulations

Different Seas, Different Phishes - Large-Scale Analysis of Phishing Simulations

ACM AsiaCCS 2025

Oskar Braun, Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann

First page of the paper: A Platform for Physiological and Behavioral Security

A Platform for Physiological and Behavioral Security

NSPW 2025

Jan Hörnemann

First page of the paper: Privacy from 5 PM to 6 AM: Tracking and Transparency in the HbbTV Ecosystem

Privacy from 5 PM to 6 AM: Tracking and Transparency in the HbbTV Ecosystem

IEEE/IFIP DSN 2025

Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann

First page of the paper: Understanding Regional Filter Lists: Efficacy and Impact

Understanding Regional Filter Lists: Efficacy and Impact

PoPETS 2025

Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann

Who is AWARE7 the right partner for?

Mid-sized companies with 50-2,000 employees

Companies that need real security, without paying for a DAX-corporation provider. Fixed price, clear scope, one point of contact.

IT managers & CISOs

Who have to argue convincingly in-house and need a report in boardroom language for that, not just technical findings.

Regulated industries

Critical infrastructure, healthcare, financial services: NIS-2, ISO 27001, DORA. We know the requirements and deliver evidence that auditors accept.

Frequently asked questions about AI penetration testing

Everything you should know before your first conversation.

An AI penetration test is an authorized security assessment of AI systems conducted by specialist experts. We simulate real attacks against your Large Language Models (LLMs), ML models, RAG systems, and AI agents - from prompt injection and jailbreaking to data exfiltration and model theft. Unlike traditional pentests, we test not just the infrastructure but the AI logic itself: guardrails, alignment, training integrity, and agent behavior.
We test all mainstream AI architectures: LLM-based chatbots and copilots (GPT, Claude, Llama, Mistral), RAG systems (Retrieval-Augmented Generation), AI agents with tool access, classical ML models (fraud detection, scoring, diagnostics), multimodal systems (image + text), and the underlying infrastructure (APIs, MLOps pipelines, vector databases). Whether self-hosted or cloud API - the testing approach is individually tailored to your architecture.
An AI pentest systematically tests your system against known vulnerability classes (OWASP Top 10 for LLMs, MITRE ATLAS). You receive a prioritized list of all findings with reproduction steps. AI red teaming goes further: we simulate creative, realistic attack scenarios over several weeks - including ones not yet covered by any taxonomy. The goal is not just a vulnerability list, but the answer: how far can a motivated attacker get against your AI-powered processes?
Prompt injection is currently the most critical vulnerability in LLM applications (OWASP LLM01). An attacker manipulates input so that the model ignores its system instructions and instead executes the attacker's commands. In direct prompt injection, this occurs via user input; in indirect prompt injection, through poisoned documents or data sources processed by the model (particularly critical in RAG systems). Consequences range from data leakage and reputational damage to remote code execution when the LLM is connected to tools or APIs.
Article 15 of the EU AI Act requires high-risk AI systems to demonstrate "an appropriate level of accuracy, robustness and cybersecurity" throughout their lifecycle - including resilience against data poisoning, adversarial attacks, and model manipulation. An AI penetration test provides exactly this evidence. For GPAI model providers, governance obligations have applied since August 2025. Our report is designed as an auditable compliance record and maps all findings to the relevant EU AI Act articles.
The OWASP Top 10 for Large Language Model Applications is the international community standard for LLM security, developed by a global community of hundreds of experts. The ten categories of the current version (v2025) include: Prompt Injection, Sensitive Information Disclosure, Supply Chain, Data and Model Poisoning, Improper Output Handling, Excessive Agency, System Prompt Leakage, Vector and Embedding Weaknesses, Misinformation, and Unbounded Consumption. We use this taxonomy as the methodological basis for every LLM pentest, supplemented by the MITRE ATLAS framework for threat modeling.
MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) is the AI-specific extension of the well-known MITRE ATT&CK framework. It documents tactics, techniques, and procedures (TTPs) from real attacks against AI systems - from reconnaissance through model evasion to data exfiltration. We use ATLAS for threat modeling your AI system and structure our red team scenarios along this attack matrix.
Our process covers five phases: 1) Scoping Workshop - identification of all AI components, threat modeling per MITRE ATLAS, definition of rules of engagement. 2) Reconnaissance - analysis of AI architecture, model endpoints, data pipelines, guardrails, and integrations. 3) Vulnerability Testing - automated scans (Garak, Promptfoo) combined with manual expert analysis for prompt injection, jailbreaking, data exfiltration, guardrail bypass, and agent behavior. 4) Exploitation - confirmation of critical findings with proof-of-concept, chaining into realistic attack scenarios. 5) Reporting - technical report with CVSS scoring, compliance mapping (OWASP, EU AI Act, ISO 42001) and prioritized remediation roadmap.
Costs depend on scope and complexity. A focused LLM pentest (single chatbot/copilot, OWASP Top 10 LLM) starts from EUR 8,100 net. A comprehensive AI security assessment covering multiple models, RAG system, and agent testing starts from EUR 14,850 net. Full AI red teaming over 4-6 weeks starts from EUR 25,650 net. You receive a binding fixed-price quote within 48 business hours - no hourly rates, no additional charges.
ISO/IEC 42001 is the international standard for AI management systems - comparable to ISO 27001 for information security, but specific to AI. The standard defines 38 controls in 9 objective categories and enables certification. For organizations deploying AI in regulated sectors (finance, healthcare, critical infrastructure), ISO 42001 is increasingly becoming a differentiator with customers and regulators. An AI pentest provides the technical evidence you need for the controls in ISO 42001.
Yes. We systematically test all protective layers of your AI application: content filters, jailbreak detectors, PII masking, output validators, and constitutional classifiers. We assess both bypass resistance (false-negative rate under adversarial conditions) and the false-positive rate (does the guardrail block legitimate usage?). You receive a quantitative evaluation of guardrail effectiveness with concrete hardening recommendations.
AI systems require more frequent testing than traditional software: models are regularly fine-tuned, RAG content changes daily, agents gain new capabilities - every change can introduce new vulnerabilities without a single line of code being modified. We recommend at minimum one full AI pentest annually, semi-annually for critical systems. For organizations with continuous model update cycles, we offer a retainer model with quarterly tests.

How secure is your AI really?

Our experts test your LLMs, RAG systems, and AI agents - with a fixed-price commitment and audit-ready reporting.

Kostenlos · 30 Minuten · Unverbindlich

Rufen Sie uns an

Mo-Fr, 8:00-17:00 Uhr - persönlich und unverbindlich.

0209 8830 6764
Jetzt anrufen