AI Red Teaming
We attack your AI in Japanese and English — on your hardware.
Your LLM was safety-trained in English. Your attackers will not be. Published research is consistent on this point: safety alignment skews heavily toward English, and adversarial prompts in other languages — plus code-switching between them — routinely defeat models that hold the line in English. Nearly every red-teaming tool on the market ships an English-only corpus. That is the gap we test.
Our harness carries a Japanese attack corpus built for how the language actually breaks models: keigo authority escalation (honorific register used to manufacture a chain of command the model obeys), full-width and kanji-variant character obfuscation that slips past string-matching filters, and My Number / PII regurgitation probes written to APPI's definition of personal information — not a translated English test.
And we run it inside your perimeter. Commercial AI-security platforms are SaaS: your production prompts, your system prompt, and your model's responses travel to a vendor's cloud, and self-hosting is rarely on the menu. For a Japanese enterprise that turns a security test into a cross-border transfer of personal data — a consent-and-disclosure problem under APPI, not a security one. We ship the harness to your network, execute against your endpoint, and hand you an evidence bundle showing a third-party API call count of zero.
What we deliver
LLM Red Team Assessment
Full adversarial pass mapped to the OWASP LLM Top 10 — prompt injection, sensitive information disclosure, excessive agency, hidden context exposure. EN + JA corpus. $12K-18K.
Japanese-Language Adversarial Testing
日本語レッドチーム. Keigo authority escalation, full-width and kanji-variant obfuscation, My Number and APPI-scoped PII regurgitation. Native corpus, not translated. Add-on or standalone, $4K-8K.
On-Premise & Air-Gapped Execution
The harness runs inside your network or VPC against your endpoint. Zero third-party API calls, evidenced in the report. Supports an APPI cross-border transfer assessment.
Hit Rates, Not Single-Shot Verdicts
Every probe is fired five times. Models sample their output, so one clean answer is a coin flip, not evidence. Two bypasses in five is a finding — an attacker retries.
Agentic & Tool-Use Boundary Testing
For AI that can act, not just answer. We test whether the model claims or attempts capability outside its authorized tool scope — the failure mode that turns a chatbot into an incident.
Continuous Regression Testing
Model upgrades silently undo safety fixes. A hash-pinned corpus re-runs on every model or prompt change and gates your CI. Retainer from $2.5K/mo.
Auditor Evidence Package
Findings mapped to OWASP LLM Top 10, NIST AI RMF, EU AI Act adversarial-testing obligations, and the METI/MIC AI Guidelines for Business. Bilingual report.
How an engagement unfolds
- 01
Scope & Authorization
A signed scope file fixes the target identifier, the authorization reference, and the test window. The harness refuses to fire at anything outside it. No ambiguity about what was authorized, ever.
- 02
Corpus Build
We tailor the probe set to your system prompt, your tool surface, and your data classes — then write the Japanese probes natively against your domain vocabulary.
- 03
On-Premise Execution
The harness runs inside your perimeter against your endpoint. Every probe and every response stays on your infrastructure. The report states the third-party call count.
- 04
Findings & Triage
Each finding carries a severity, an OWASP mapping, the exact prompt, a bypass count out of five attempts, and the response that proves it. We seed canary tokens out of band so a leak is demonstrated, not inferred from a model repeating your words back.
- 05
Re-Test & Regression Gate
The corpus is SHA-256 pinned, so after you remediate the identical test set re-runs and we compare hit rates before and after. A probe that went from four bypasses in five to zero is a fix; one that went from four to three is noise. That comparison then stays in your pipeline as a gate.
