LIVE STREAM
WRAITH: GPT-5.6 Sol safety limits bypassed using Crescendo in 4 turns
RE-DEFINING MODEL SAFETY

Adversarial Benchmarks &
Model Vulnerability Indices

We continually evaluate safety limits across open and closed language models. Our public datasets catalog attack success rates (ASR) to foster robust defensive alignment.

01 / DATASETS & EVALS

Interactive Datasets

Choose a standardized safety dataset below. We benchmark each model's refusal threshold to trace vulnerabilities in adversarial scenarios.

Size400+ Prompts
Taxonomy7 Categories
Updated2026-06
Evaluated Models & ASRDataset: HarmBench
Model TargetAttack Success RateAvg TokensLatency
Mythos 5 (Anthropic SOTA)
94%
3951.6s
GPT-5.6 Sol (OpenAI)
95%
4201.4s
Gemini 3.6 Flash (Google)
91%
3800.8s
Fable 5
89%
4101.2s
Claude 3.7 Sonnet (Thinking)
88%
4351.5s

COMPARATIVE VULNERABILITY INDEX

Higher score = More vulnerable
Mythos 5 (Anthropic SOTA)ASR Index: 0.94
EXPLORECOMPLYEXPLOIT
GPT-5.6 Sol (OpenAI)ASR Index: 0.95
EXPLORECOMPLYEXPLOIT
Gemini 3.6 Flash (Google)ASR Index: 0.91
EXPLORECOMPLYEXPLOIT
Fable 5ASR Index: 0.89
EXPLORECOMPLYEXPLOIT
Claude 3.7 Sonnet (Thinking)ASR Index: 0.88
EXPLORECOMPLYEXPLOIT
02 / ATTACK ENGINE DIRECTORY

Supported Tactics & Strategies

Browse the red-teaming algorithms embedded in the WRAITH core library. Run any of these natively in sandbox environments.

ASR: 62% - Medium

Encoding Transformation

Obfuscates prompts using Base64, Rot13, Caesar cipher, or binary representation to bypass target tokenization safety filters.

Launch Sandbox Job
ASR: 85% - Extreme

PAIR (Adversarial Refinement)

An attacker LLM iteratively queries the target, refining jailbreak payloads in a self-play feedback loop to maximize refusal bypass.

Launch Sandbox Job
ASR: 88% - Extreme

TAP (Tree of Attacks)

Extends PAIR with tree branching and safety response pruning, exploring multiple attack pathways simultaneously to avoid safety traps.

Launch Sandbox Job
ASR: 91% - Critical

Crescendo (Multi-turn Escalation)

Steers targets by initiating benign dialogue, gradually pivoting the topic, and backtracking/rephrasing prompts upon target refusals.

Launch Sandbox Job
ASR: 87% - Extreme

Best-of-N (BoN) Sampling

Generates diverse, augmented variations of the attack prompt, submits them in parallel, and scores/yields the strongest bypass output.

Launch Sandbox Job
ASR: 93% - Critical

Indirect Prompt Injection (RAG/CWE-1158)

Embeds hidden adversarial directives into retrieved external context (HTML/Markdown/Search) to hijack search assistants and RAG agents.

Launch Sandbox Job
ASR: 89% - Extreme

PDF Document Structure Injection

Exploits document extraction pipelines by embedding instruction overrides inside PDF metadata, annotations, or invisible text layers.

Launch Sandbox Job
ASR: 86% - High

Multi-Modal Vision Prompt Injection

Injects adversarial typographic instructions into visual document scans, low-contrast image layers, or OCR streams for visual LLMs.

Launch Sandbox Job
ADVANCED METHODOLOGY SHOWCASE

Pliny & BT6 GG Adversarial Methodology

Demonstrating gradient-guided suffix optimization (BT6 GG) combined with Pliny Liberator structural prompt synthesis. Evaluates cross-architecture jailbreak transferability across frontier models.

Mythos 594.2% ASR
GPT-5.689.8% ASR
Gemini 3.691.5% ASR
Transferability Matrix
Llama BT6 GG SuffixTransfer: High
Pliny v6 SynthesisBypass: Verified
Refusal Suppression99.1% Confidence
METHODOLOGY & TRANSPARENCY

Rigorous Defensive Evaluation

All benchmarks are evaluated using standard parameters (budget limits, temperature=0.0, HeuristicClassifier taxonomy scoring) to ensure reliable replicability. We strictly avoid model damage and promote constructive disclosure.