The guardrail dilemma
Too restrictive leads to user frustration and guardrail deactivation. Too tolerant lets attackers through. Correct calibration requires systematic testing, not intuition.
Guardrail Assessment
Content filters, jailbreak detectors, PII masking - your guardrails only protect your AI system as well as they perform against real attacks. We measure what they actually deliver: quantitatively, reproducibly, audit-ready.
The Problem
Most AI guardrails are configured, not tested. They are switched on, the out-of-the-box defaults adopted and considered "secure". The reality: jailbreak techniques evolve daily. What was blocked yesterday gets through today with a minimal reformulation. And nobody measures it.
The guardrail dilemma
Too restrictive leads to user frustration and guardrail deactivation. Too tolerant lets attackers through. Correct calibration requires systematic testing, not intuition.
Jailbreaking is industrialised
Public databases with thousands of jailbreak prompts, automated bypass tools and community forums make it trivially easy to circumvent unpatched guardrails.
Regulatory requirements
EU AI Act Art. 15 and the GPAI Code of Practice require demonstrable robustness against misuse. A Guardrail Assessment provides measurable evidence of this robustness.
Latency as an attack surface
External guardrail services add latency. Under adversarial load - complex requests deliberately saturating the classifiers - response time can rise to several seconds and cause timeouts.
GUARDRAIL LAYERS - WHAT MUST BE TESTED
Input Classifier
Detects malicious inputs before LLM processing
System Prompt Guards
Protects system prompt from being overwritten by user inputs
Output Classifier
Filters harmful outputs after LLM processing
PII Masking
Anonymises personal data in outputs
Topic Restriction
Limits discussion scope to permitted topics
Constitutional Classifier
Checks outputs against ethics guidelines and policies
Monitoring & Logging
Detects bypass patterns in real time
FROM OUR PRACTICE
Significant bypass rates
Our assessments regularly reveal that a significant proportion of tested guardrails can be bypassed - particularly with uncalibrated out-of-the-box deployments.
After hardening based on our recommendations, bypass rates can be substantially reduced while maintaining an acceptable false-positive rate.
What we test
The Guardrail Assessment measures all relevant quality dimensions of your AI protection layers - quantitatively and comparably.
What percentage of all adversarial requests are correctly blocked? We test with 500+ curated bypass techniques: roleplay prompts, token smuggling, adversarial suffixes, multilingual exploits, many-shot jailbreaking, contextual circumvention and novel, non-publicly known techniques from our research.
False-Negative Rate
500+ test cases
A guardrail that blocks every other legitimate request is not protection - it is a productivity killer. We measure the false-positive rate with realistic, harmless requests from your use case and identify the optimal calibration point between security and usability.
Usability Impact
Usability Balance
ROC Curve
Attackers can saturate guardrails with deliberately complex requests - a timing attack on your AI application. We measure latency under normal operation and under adversarial load: P50, P95, P99 response times and timeout behaviour.
Performance Degradation
P95 Latency
Timeout Resistance
A guardrail configured for GPT-4 may be ineffective for Claude or Llama - each model has different tokenisation properties and behaviour patterns. We test whether your guardrails work model-independently or need model-specific calibration.
Portability
Multi-Model
Tokeniser Differences
Many guardrails only check individual messages, not the conversation history. Attackers use multi-turn sequences to gradually condition guardrails: a harmless request sets up the next one until the classifier is outwitted. We test robustness over 10-, 20- and 50-turn conversations.
Conversation Robustness
Multi-Turn
Context Attacks
Does your system disclose personal data despite guardrails - from context, from training or from connected data sources? We test PII masking effectiveness, membership inference resistance and contextual data exfiltration per GDPR requirements.
Privacy Compliance
GDPR
PII Masking
Tested Systems
Every guardrail system has its own vulnerability classes and bypass techniques - generic tests are not sufficient.
Microsoft
Specific bypass techniques for Azure AI Content Safety: severity threshold exploits, category-specific circumventions (Hate, Violence, Sexual, Self-Harm), Prompt Shield bypass and multi-modal attacks. We test all four harm categories and groundedness detection.
AWS
Testing of all Bedrock guardrail functions: topic denial, content filter, word filter, sensitive information filter and grounding checks. Specific attacks on topic restriction bypasses and PII entity recognition gaps in non-English text.
NVIDIA
Colang-based guardrail flows have specific logic exploits: flow manipulation through adversarial inputs, rail bypass via uncovered conversation paths and input/output rail inconsistencies. We test both predefined and custom flows.
Anthropic
Training-based guardrails have different vulnerabilities than external classifiers: contextual conditioning, reasoning exploits, cross-lingual bypasses and adversarial roleplaying techniques that circumvent Constitutional AI checks. We test Claude models in their production context.
Open Source / SaaS
Real-time guardrail API with specialised prompt injection detectors: we test detection rates against current jailbreak databases, latency under load and effectiveness for non-English language inputs that may be underrepresented in the training set.
Proprietary
Many organisations build their own guardrails based on regex, keyword lists or fine-tuned classifiers. We analyse your specific implementation, identify gaps in coverage and develop bespoke bypass tests and hardening recommendations.
Your deliverable
The Guardrail Effectiveness Score (GES) is the central deliverable of the assessment. It compresses the security performance of your guardrails into an understandable, comparable metric - as the basis for management decisions and compliance evidence.
Guardrails are effectively calibrated. Bypass rate < 5%, false-positive rate < 8%. Targeted optimisation recommended.
Basic protection in place, but bypass rate 5-15% or elevated false-positive rate. Specific hardening measures recommended.
Significant vulnerabilities. Bypass rate 15-30%. Structural reconfiguration required.
Guardrails provide no reliable protection. Bypass rate > 30%. Immediate action required before production use.
COMPONENTS OF THE GUARDRAIL EFFECTIVENESS SCORE
COMPLIANCE USE OF THE GES
Methodology
Quantitative testing with 500+ curated test cases - combined with manual expert analysis for novel bypass techniques.
Complete mapping of all guardrail layers: which systems are active? How are they configured? Which harm categories are covered? Which thresholds are set? Which models do they run on? Result: guardrail architecture diagram.
1 day
Establishing the initial measurement: false-positive rate with 200+ legitimate requests from your use case. Latency baseline under normal operation. Performance profile of the guardrail system as reference for all subsequent tests.
1-2 days
Testing with 500+ curated bypass techniques from our proprietary test set - broken down by attack category: jailbreaking, roleplays, token smuggling, encoding tricks, multilingual exploits, many-shot conditioning and adversarial suffixes. False-negative rate measured per category.
3-5 days
Tests not covered by single-message analysis: stepwise context conditioning over 10-, 20-, 50-turn conversations. Guardrail exhaustion attacks. Cross-session persistence tests for systems with persistent conversation memory.
2-3 days
Quantitative latency measurement under adversarial load: P50, P95, P99 response times. Timeout behaviour with complex requests. Guardrail system behaviour under overload - does it fail open or fail closed?
1-2 days
Guardrail Effectiveness Score (GES) with breakdown by test dimension. Concrete calibration recommendations for each threshold. Hardening roadmap with prioritisation. Compliance mapping to EU AI Act, ISO 42001 and GDPR.
2-3 days
Typical total duration: 8-15 days - depending on the number of guardrail layers and desired test depth.
You receive a binding fixed-price quote within 48 business hours from EUR 10,000.
Why AWARE7
Pure awareness platforms don't test systems. Pure consulting firms are too far removed. AWARE7 combines both: we hack your infrastructure and train your employees: tailored to mid-sized companies, personal, without enterprise overhead.
Around 20% of our revenue comes from research projects for the BSI and the BMBF. Our studies, published at ACM and Springer conferences, analyse millions of websites and tens of thousands of phishing emails. Three of our executives are professors at German universities at the same time.
From first contact to final report, your data is stored on our own servers in Germany - no US cloud providers, no third-country transfers. Our AI also runs on our own hardware in Germany - with locally operated open-source models. Client and project data never reach external AI services. All staff are permanently employed, covered by social insurance and bound by uniform legal obligations.
More on digital sovereigntyWithin 24 hours you receive a binding fixed-price quote without hourly rate risk. A well-practised team and standardised processes ensure a clear schedule with a defined start and end date.
A personal project manager accompanies you from the first meeting to the retest. You book appointments directly with your contact person and keep the same contact throughout the project.
Peer-reviewed publications
Different Seas, Different Phishes - Large-Scale Analysis of Phishing Simulations
ACM AsiaCCS 2025
Oskar Braun, Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann
A Platform for Physiological and Behavioral Security
NSPW 2025
Jan Hörnemann
Privacy from 5 PM to 6 AM: Tracking and Transparency in the HbbTV Ecosystem
IEEE/IFIP DSN 2025
Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann
Understanding Regional Filter Lists: Efficacy and Impact
PoPETS 2025
Jan Hörnemann, Norbert Pohlmann, Matteo Große-Kampmann
Who is AWARE7 the right partner for?
Mid-sized companies with 50-2,000 employees
Companies that need real security, without paying for a DAX-corporation provider. Fixed price, clear scope, one point of contact.
IT managers & CISOs
Who have to argue convincingly in-house and need a report in boardroom language for that, not just technical findings.
Regulated industries
Critical infrastructure, healthcare, financial services: NIS-2, ISO 27001, DORA. We know the requirements and deliver evidence that auditors accept.
Everything about guardrail bypasses, false-positive rates and the Guardrail Effectiveness Score.
We measure the effectiveness of your AI safety filters quantitatively - with 500+ bypass techniques and the Guardrail Effectiveness Score. Fixed-price commitment from EUR 10,000.
Free · 30 minutes · No obligation
Arturs Nikitins
Initial consultation & needs analysis
Looking for personal advice?
No obligation · Reply within 24h on business days