L
AI Red Team

AI Red Teaming

We attack your AI in Japanese and English — on your hardware.

Your LLM was safety-trained in English. Your attackers will not be. Published research is consistent on this point: safety alignment skews heavily toward English, and adversarial prompts in other languages — plus code-switching between them — routinely defeat models that hold the line in English. Nearly every red-teaming tool on the market ships an English-only corpus. That is the gap we test.

Our harness carries a Japanese attack corpus built for how the language actually breaks models: keigo authority escalation (honorific register used to manufacture a chain of command the model obeys), full-width and kanji-variant character obfuscation that slips past string-matching filters, and My Number / PII regurgitation probes written to APPI's definition of personal information — not a translated English test.

And we run it inside your perimeter. Commercial AI-security platforms are SaaS: your production prompts, your system prompt, and your model's responses travel to a vendor's cloud, and self-hosting is rarely on the menu. For a Japanese enterprise that turns a security test into a cross-border transfer of personal data — a consent-and-disclosure problem under APPI, not a security one. We ship the harness to your network, execute against your endpoint, and hand you an evidence bundle showing a third-party API call count of zero.

Offerings

What we deliver

LLM Red Team Assessment

Full adversarial pass mapped to the OWASP LLM Top 10 — prompt injection, sensitive information disclosure, excessive agency, hidden context exposure. EN + JA corpus. $12K-18K.

Japanese-Language Adversarial Testing

日本語レッドチーム. Keigo authority escalation, full-width and kanji-variant obfuscation, My Number and APPI-scoped PII regurgitation. Native corpus, not translated. Add-on or standalone, $4K-8K.

On-Premise & Air-Gapped Execution

The harness runs inside your network or VPC against your endpoint. Zero third-party API calls, evidenced in the report. Supports an APPI cross-border transfer assessment.

Hit Rates, Not Single-Shot Verdicts

Every probe is fired five times. Models sample their output, so one clean answer is a coin flip, not evidence. Two bypasses in five is a finding — an attacker retries.

Agentic & Tool-Use Boundary Testing

For AI that can act, not just answer. We test whether the model claims or attempts capability outside its authorized tool scope — the failure mode that turns a chatbot into an incident.

Continuous Regression Testing

Model upgrades silently undo safety fixes. A hash-pinned corpus re-runs on every model or prompt change and gates your CI. Retainer from $2.5K/mo.

Auditor Evidence Package

Findings mapped to OWASP LLM Top 10, NIST AI RMF, EU AI Act adversarial-testing obligations, and the METI/MIC AI Guidelines for Business. Bilingual report.

Process

How an engagement unfolds

  1. 01

    Scope & Authorization

    A signed scope file fixes the target identifier, the authorization reference, and the test window. The harness refuses to fire at anything outside it. No ambiguity about what was authorized, ever.

  2. 02

    Corpus Build

    We tailor the probe set to your system prompt, your tool surface, and your data classes — then write the Japanese probes natively against your domain vocabulary.

  3. 03

    On-Premise Execution

    The harness runs inside your perimeter against your endpoint. Every probe and every response stays on your infrastructure. The report states the third-party call count.

  4. 04

    Findings & Triage

    Each finding carries a severity, an OWASP mapping, the exact prompt, a bypass count out of five attempts, and the response that proves it. We seed canary tokens out of band so a leak is demonstrated, not inferred from a model repeating your words back.

  5. 05

    Re-Test & Regression Gate

    The corpus is SHA-256 pinned, so after you remediate the identical test set re-runs and we compare hit rates before and after. A probe that went from four bypasses in five to zero is a fix; one that went from four to three is noise. That comparison then stays in your pipeline as a gate.

Proof
0 bytes
Prompt or response data leaving your network
Questions

Frequently asked

Where does our data go during a test?
Nowhere. The harness executes on hardware you control, against an endpoint you control. The evidence bundle records the inference endpoint and a third-party API call count — zero on an on-premise engagement. That line exists specifically so your privacy officer can close a cross-border transfer question with a document instead of a vendor's assurance.
Why is Japanese-language testing a separate line item? Can't you translate the English tests?
Translation misses the attacks that matter. Keigo authority escalation only works in a language with grammaticalized honorific register — there is no English equivalent to translate. Full-width and kanji-variant obfuscation exploits the Japanese character set itself. And My Number has a specific format and a specific legal status under APPI that a translated 'SSN' probe does not test. The corpus is written in Japanese, for Japanese failure modes.
How is this different from an AI-security SaaS platform?
Those platforms license per AI asset on an annual contract, run English-first corpora, and process your prompts in their cloud. That is a reasonable fit for a US-only product team. It is a poor fit if your data cannot leave the country, if your users speak Japanese, or if you need a point-in-time attested report rather than a dashboard. We are engagement-priced, on-premise, and bilingual.
Do you need our model weights or our API keys?
No. We need a callable endpoint and a scope authorization. We test the system as an attacker reaches it — the model plus your system prompt plus your tools plus your guardrails — because that composite is what actually fails, not the base model in isolation.
Is the assessment reproducible?
The test set is, and we are precise about the part that is not. The report publishes a SHA-256 hash of the probe corpus, so the identical questions can be put to your system again — a changed hash means the test set moved and any comparison is invalid. Target responses are a different matter: language models sample, so the same probe against an unchanged system genuinely produces different answers run to run. That is why we fire each probe five times and report a hit rate. On re-test you compare hit rates, not a results hash. Any vendor promising byte-identical findings from a stochastic system is describing something their tool does not do.
Why five attempts per probe instead of one?
Because one attempt measures luck. A guardrail that holds four times in five is not a guardrail — an attacker simply sends the request again — but a single-shot harness that happened to sample the fourth attempt would report it as a pass. Conversely a single leak, seen once, could be an outlier worth less alarm than a probe that bypasses every time. Reporting two-in-five versus five-in-five tells your engineers which fix is urgent. We will run more than five attempts on request where the stakes justify the inference cost.
What does an engagement cost and how long does it take?
A standard bilingual assessment is $12K-18K and runs two to three weeks from signed scope to delivered report. For reference, the market range for comparable one-time AI red team engagements is roughly $8K-25K, with large multi-agent programs reaching $50K-150K. Continuous regression retainers start at $2.5K/mo.

Ready to get started?

Book a consultation or send us the scope of what you need.

Book a Consultation