Skip to content
CloudSecOps

AI red teaming

We run objective-driven adversarial campaigns against your LLM and agent deployments — multi-turn jailbreak and goal-hijack chains, scored as a measured bypass rate per control layer rather than a pass/fail verdict.

When you need this

  • You're weeks from launching a customer-facing assistant and someone senior asked whether anyone has genuinely tried to break it — not scanned it, broken it.

  • Your guardrails block the obvious prompts, so the question has moved on: nobody in the room can put a number on how much of a determined, multi-turn attempt still gets through.

  • The underlying model changed — a version bump, a cheaper tier, a new fine-tune — and nobody can say whether the refusal behaviour you validated last quarter survived it.

  • Your governance or GRC team needs adversarial-testing evidence for an AI risk register, a customer security review, or an ISO/IEC 42001 audit, and screenshots of a chat session aren't it.

What we do

01

Scenarios written against your business objectives

We start from what an attacker would want out of your deployment, not a generic harm checklist. That means reading your system prompts, tool schemas, tenancy model, and the policies the assistant is meant to enforce, then defining concrete adversary objectives: issue a refund outside policy, surface another tenant's records through retrieval, get regulated advice published under your brand voice, escalate an agent's own permissions. Every campaign is scored against those objectives, so a result means something to the business and not only to the model.

02

Campaigns, not single prompts

Single-shot payloads are the part automation already handles. The attacks that land are multi-turn: gradual escalation that builds on the model's own prior answers (Crescendo-style), many-shot context stuffing, iterative refinement in the style of PAIR and TAP, persona and fictional-framing pivots, encoding and low-resource-language obfuscation, and best-of-N resampling against nondeterministic refusals. Where the system carries memory across sessions, we seed in one session and collect in another. We run garak, PyRIT, and promptfoo for breadth and regression coverage; the campaigns that break business logic are designed and driven by the engineer on the engagement.

03

Guardrail bypass measured per layer

For each objective and harm category we fix a corpus, run it at a stated sample size, and record attack success rate — then attribute each block to the layer that actually produced it: input classifier, system prompt, model refusal, output filter, or tool authorization. This separates controls that stop an attack from controls that only slow the first attempt, and it distinguishes a real refusal from a deflection that leaks the answer two turns later. Corpora, sample sizes, and scoring criteria are kept so the same measurement re-runs after your fixes and the number moves, or doesn't.

04

Objectives carried through to a real side effect

A campaign only counts when it ends in something that happened, not in a paragraph the model should never have written. Where the deployment can act, we drive each objective to its consequence and record it — the transaction that posted, the other tenant's record that came back, the outbound message that actually sent. That usually means arriving through data rather than the chat box: a payload seeded in a document the system will retrieve, or a memory written in one session and collected in a later one. Each objective is run repeatedly rather than once, because a chain that lands one attempt in twelve is a different risk from one that lands every time, and both differ from one that only lands on the fourth turn. Successful chains carry OWASP LLM Top 10 (2025) identifiers — LLM01 prompt injection, LLM02 sensitive information disclosure, LLM06 excessive agency — plus MITRE ATLAS technique IDs, and where the agent acts, the matching OWASP Agentic Top 10 entry from ASI01–ASI10. Where the root cause is structural rather than linguistic — a tool scoped wider than its task, an identity the agent should not hold, an approval step that shows a summary the agent wrote instead of the parameters it will send — we name it and route it to agent security work, because no amount of prompt tuning closes it.

What you receive

  • A scenario library — adversary objectives, harm categories, and rules of engagement agreed before testing and reusable afterwards
  • Bypass-rate tables per objective, harm category, and control layer, with the sample size and model or prompt version each number was measured against
  • Full transcripts of the campaigns that succeeded, written end to end: entry point, escalation turns, and the side effect they reached
  • Control-layer analysis that says which mitigation to change — authorization boundary, tool scope, retrieval filtering, system prompt — and which existing ones are cosmetic
  • A governance evidence pack mapping the exercise to the MEASURE function of the NIST AI RMF — principally MEASURE 2.7, evaluating and documenting AI system security and resilience — and to the ISO/IEC 42001 Annex A controls covering AI system verification and validation and AI system operation and monitoring
  • A second measurement after remediation against the identical corpus, reported as before-and-after bypass rates per layer, with the scoring rubric and harness configuration handed over so your team can produce the same number without us

How it runs

  1. Scenario design

    days

    Threat model, adversary objectives, harm categories relevant to the deployment, and testing boundaries agreed in writing.

  2. Campaign

    1–3 weeks

    Automated coverage sweeps plus manual multi-turn campaigns. Anything reaching a real-world side effect is escalated on discovery.

  3. Measure & report

    days

    Bypass rates per layer, successful chains, control recommendations, and the governance mapping.

  4. Re-run

    included

    Same corpus, same sample sizes, same scoring criteria, run after your fixes — so the two bypass rates are comparable.

Durations are indicative. Actual scope and price are fixed after the scoping call.

Questions we get asked

How is this different from an AI penetration test?
A penetration test is breadth against a system at a point in time: the surface, the auth model, the data flows, the configuration. Red teaming is depth against objectives — can a determined person get this system to do a specific thing it must not do, given as many turns as they want. The outputs differ too. A pentest hands you findings; a red team hands you a bypass rate, the chains that produced it, and which control layer is carrying the weight. Teams often run the pentest first and the red team before a launch or a major model change.
We already run an automated red-teaming product. Why pay for people?
Those tools are useful and we run them ourselves — garak, PyRIT, and promptfoo are in every engagement for coverage and regression. What they do is generate variations of known attacks against generic harm categories and score them with a classifier. What they don't know is your refund policy, your tenant boundary, or which of your tool calls is dangerous in combination with which other one. Business-logic attacks have to be invented, and scoring them requires knowing what "bad" means in your context. The arrangement that works: automated suites in CI for regression, humans periodically for the creative work.
Will this satisfy our auditor or a customer's security review?
It gives them something concrete to file: documented adversary objectives, method, sample sizes, results per control layer, and remediation status, mapped to NIST AI RMF and ISO/IEC 42001 Annex A controls. That is usually what a governance team is missing. What it is not: we are not a certification body, we don't issue attestations, and no report certifies that a system is safe. It is evidence that adversarial testing was performed on a stated version, which is the claim the frameworks actually ask you to support.
Which harm categories do you test, and are there any you won't?
We test the categories that match the deployment and its jurisdiction. An internal engineering assistant is tested for data disclosure, excessive agency, and prompt leakage; a consumer-facing brand assistant adds regulated advice, defamation, and discriminatory output in anything touching decisions about people. We don't run generic safety benchmarks for their own sake. There are limits we hold to: we do not generate or test child sexual abuse material or content whose creation is itself unlawful, and we will not assert that a model is safe against a category we did not test. For frontier-model safety evaluations of that kind, the model provider's own evaluations and specialist labs are the right route, and we'll say so rather than take the work.