Home / Services / LLM Firewall
Input inspection for generative AI appsA firewall at the entrance
to your generative AI apps.
LLM Firewall is an API service that inspects user input before it is passed to an LLM. It detects prompt injection, personal information or credentials mixed into input, web attacks such as SQL injection, and harmful content. Input hidden with Base64, leetspeak and similar techniques is decoded before it is examined. Results are returned per detector, so your app decides whether to allow or block.
Challenges
Sound familiar?
Prompt injection
Input that tries to override instructions can manipulate the AI's behavior or its internal information.
Confidential data sent out
Users may type personal information or API keys as-is, and that input gets sent to an external LLM.
Invisible to existing WAFs
Attacks blended into natural language and instructions hidden by encoding are beyond what conventional WAF patterns can catch.
Features
LLM Firewall key features
01Prompt injection detection
Combines machine learning models with high-precision rules to detect input aimed at hijacking instructions or bypassing restrictions.
02Personal and confidential information detection
Detects credit card numbers, API keys and credentials, as well as Japan-specific data such as My Number and corporate numbers, and can mask the values.
03Web attack pattern detection
Detects SQL injection, XSS, command injection, SSRF and more after restoring input to its original form through multi-stage decoding.
04Harmful content by category
Classifies harmful content by category. Enable or disable categories per project or per API key to match your use case.
05Inspection after deobfuscation
Based on joint research with a university, it restores Base64 and other encodings, ASCII art and split instructions to their original form to find hidden instructions.
06Reacts to unknown obfuscation
Detects input that matches no known technique but has an unnatural structure as “unknown obfuscation”, to prepare for new evasion techniques.
07Japanese input support
Japanese input is translated into English by a translation model we operate ourselves, then inspected. Input is not sent to external translation services.
08Policies, logs and alerts
Enable or disable detectors and personal information rules per project, and override them per API key. Includes API logs, audit logs and alerts.
Layered inspection
One API call,
multiple detectors at once
LLM input carries more than one kind of risk. LLM Firewall runs multiple detectors at once in a single API call and returns per-detector results along with an overall severity. You choose which detectors to use per project.
Injection
Combines machine learning models and rules to identify input aimed at overriding instructions or bypassing restrictions.
Personal and confidential information
Detects card numbers, API keys, Japan-specific identification numbers and more, and returns the values masked when needed.
Web attacks
Identifies attack patterns such as SQL, command and SSRF attacks that could reach your backend through the LLM.
Harmful content
Classifies harmful content by category, and lets you enable only the categories relevant to your use case.
Joint research
Joint university research,
built into the inspection engine
“Non-natural language prompt injection”, which hides instructions with encoding, ASCII art or meaningless strings, easily slips past defenses that assume natural language. In joint research by the Japan Advanced Institute of Science and Technology (JAIST) and Classmethod Security (formerly CyberMatrix), 14 major LLMs were evaluated, showing that this type of attack bypassed safety measures at rates of 38-52%. LLM Firewall implements the countermeasures for each technique shown in this research in its inspection engine, restoring hidden instructions to their original form before analysis.
Encoded instructions
Decodes and normalizes Base64, hexadecimal, confusable Unicode characters, zero-width characters, leetspeak and more before inspection.
ASCII art text
Runs characters drawn with symbols through optical character recognition (OCR) as an image, and inspects the words it reads.
Split instructions
Traces instructions split across multiple variables and reassembles them into the original continuous instruction before analysis.
Appended meaningless strings
Finds and reports meaningless symbol sequences and unnatural repetition appended after an instruction.
Unknown techniques
Detects input that matches no known technique but has an unnatural structure as “unknown obfuscation”, to prepare for new evasion techniques.
Japanese, too
Japanese input is translated by a translation model we operate ourselves before inspection, so attacks written in Japanese are covered as well.
Paper:Experimental Evaluation of Non-Natural Language Prompt Injection Attacks on LLMs (SPAIML'25 international workshop, October 2025, CEUR Workshop Proceedings Vol-4154)
Cloud-agnostic
Wherever your LLM runs,
the same standard of protection
LLM Firewall is an API that does not depend on any particular cloud or LLM. Whether you use a cloud LLM such as Amazon Bedrock, Azure OpenAI or Google Gemini, an open LLM on your own servers, or an on-premises business system, you get the same inspection just by calling it over HTTPS.
No cloud account needed
All you need is a CyberForces API key. There is no need to set up AWS, Azure or Google Cloud accounts or permissions.
On-premises and self-hosted LLMs
Business apps in your own data center and LLMs you operate yourself can use it, as long as they can make outbound HTTPS connections.
One standard for multiple LLMs
Even if you use different LLMs for different purposes, you manage input inspection rules in a single policy.
Settings survive an LLM switch
When you switch models or clouds, there is no need to rebuild your detector, personal information rule or alert settings.
Please note: LLM Firewall is a cloud service (API) operated by CyberForces. It is not deployed in your environment. Use requires outbound HTTPS connectivity to the inspection API (Tokyo Region).
LLM Firewall × Cloud guardrails
Use it alongside cloud guardrails,
with separate roles
Amazon Bedrock Guardrails, Azure AI Content Safety and Google Cloud Model Armor are strong at controlling harmful content and at checking model output and the grounding of responses. LLM Firewall specializes in detecting security attacks hidden in input, especially web attacks such as SQL injection and SSRF. Whichever cloud's LLM you use, you can combine them: LLM Firewall on the input side and each provider's guardrails on the output side.
| LLM Firewall | Amazon Bedrock Guardrails | Azure AI Content Safety | Google Cloud Model Armor | |
|---|---|---|---|---|
| Primary role | Detecting security attacks hidden in input | Controlling content safety and response quality | Detecting harmful content and protecting prompts | Safety filters for input and output |
| What is inspected | Input to the LLM | Input and output | Input and output | Input and output |
| Prompt injection | ◯ Combination of ML and rules | ◯ Prompt attack filter | ◯ Prompt Shields | ◯ Prompt injection and jailbreak detection |
| Web attacks (SQL, XSS, SSRF, etc.) | ◯ Analyzed after multi-stage decoding | — | — | — Malicious URL detection supported |
| Japan-specific identifiers (My Number, etc.) | ◯ Detected out of the box | △ You define regular expressions yourself | ◯ PII detection in Azure AI Language (separate service) | ◯ Integration with Sensitive Data Protection |
| Harmful content | ◯ Selectable by category | ◯ | ◯ | ◯ |
| Denied topics / banned words | — | ◯ Denied topics and word filters | ◯ Blocklists and custom categories | — |
| Output quality (grounding checks, etc.) | — | ◯ Grounding and Automated Reasoning | ◯ Groundedness detection (preview) and protected material detection | — |
| Management | CyberForces console (audit logs, alerts) | AWS Management Console / API | Azure portal / Content Safety Studio / API | Google Cloud console / API |
Benefits of using both
Two layers, different mechanisms
Stacking two inspections with different detection mechanisms raises the chance that input slipping past one is stopped by the other.
One standard on the input side
Even across multiple clouds and LLMs, input inspection is unified in LLM Firewall, with no effort spent aligning settings.
Separate roles, separate owners
The security team manages input-side security centrally, while each app team tunes output quality with guardrails.
Same screens as your other defenses
Detections of attacks on your LLMs land in the same CyberForces console, audit logs and alerts as WAAP and CSEM.
Attacks never reach the model
Input stopped at the entrance is never passed to the LLM, so you pay no inference cost for it.
One price for all inspections
Injection, personal information, web attacks, harmful content and obfuscation are inspected for a single price based on the number of tokens processed.
Please note: Capabilities of other providers' services are described based on each provider's public documentation as of October 2026. Check each provider's documentation for current support. Amazon Bedrock is a trademark of Amazon.com, Inc. or its affiliates, Azure is a trademark of Microsoft Corporation, and Google Cloud is a trademark of Google LLC.
Specifications
Scope and delivery
| Integration |
|
|---|---|
| Verdict |
|
| Detectors |
|
| Management |
|
| Pricing |
|
Please note: Inspection covers input text sent to the LLM. Inspection of LLM output and deployment as a transparent proxy in the network path are not supported.
Getting started
Getting started
Create a project
Create a project and issue an API key in the console.
Integrate the API
Add a call to the inspection API just before your app calls the LLM.
Tune the policy
Review the logs and adjust detectors, personal information rules and harmful content categories to fit your use case.
FAQ
FAQ
Is input blocked automatically when something is detected?
LLM Firewall is an API that returns a verdict. Your app decides whether to block or just warn.
Is the text I submit sent to outside parties?
Inspection runs within the CyberForces environment, and no external translation service is used for Japanese translation either.
Works well with
LLM Firewall — details and demo requests
Our team will explain deployment options and pricing for your environment.