Home / Services / LLM Firewall

Input inspection for generative AI apps

A firewall at the entrance
to your generative AI apps.

LLM Firewall is an API service that inspects user input before it is passed to an LLM. It detects prompt injection, personal information or credentials mixed into input, web attacks such as SQL injection, and harmful content. Input hidden with Base64, leetspeak and similar techniques is decoded before it is examined. Results are returned per detector, so your app decides whether to allow or block.

Challenges

Sound familiar?

Prompt injection

Input that tries to override instructions can manipulate the AI's behavior or its internal information.

Confidential data sent out

Users may type personal information or API keys as-is, and that input gets sent to an external LLM.

Invisible to existing WAFs

Attacks blended into natural language and instructions hidden by encoding are beyond what conventional WAF patterns can catch.

Features

LLM Firewall key features

01Prompt injection detection

Combines machine learning models with high-precision rules to detect input aimed at hijacking instructions or bypassing restrictions.

02Personal and confidential information detection

Detects credit card numbers, API keys and credentials, as well as Japan-specific data such as My Number and corporate numbers, and can mask the values.

03Web attack pattern detection

Detects SQL injection, XSS, command injection, SSRF and more after restoring input to its original form through multi-stage decoding.

04Harmful content by category

Classifies harmful content by category. Enable or disable categories per project or per API key to match your use case.

05Inspection after deobfuscation

Based on joint research with a university, it restores Base64 and other encodings, ASCII art and split instructions to their original form to find hidden instructions.

06Reacts to unknown obfuscation

Detects input that matches no known technique but has an unnatural structure as “unknown obfuscation”, to prepare for new evasion techniques.

07Japanese input support

Japanese input is translated into English by a translation model we operate ourselves, then inspected. Input is not sent to external translation services.

08Policies, logs and alerts

Enable or disable detectors and personal information rules per project, and override them per API key. Includes API logs, audit logs and alerts.

Layered inspection

One API call,
multiple detectors at once

LLM input carries more than one kind of risk. LLM Firewall runs multiple detectors at once in a single API call and returns per-detector results along with an overall severity. You choose which detectors to use per project.

INJECTION

Injection

Combines machine learning models and rules to identify input aimed at overriding instructions or bypassing restrictions.

PII / SECRETS

Personal and confidential information

Detects card numbers, API keys, Japan-specific identification numbers and more, and returns the values masked when needed.

WEB ATTACK

Web attacks

Identifies attack patterns such as SQL, command and SSRF attacks that could reach your backend through the LLM.

HARMFUL

Harmful content

Classifies harmful content by category, and lets you enable only the categories relevant to your use case.

Joint research

Joint university research,
built into the inspection engine

“Non-natural language prompt injection”, which hides instructions with encoding, ASCII art or meaningless strings, easily slips past defenses that assume natural language. In joint research by the Japan Advanced Institute of Science and Technology (JAIST) and Classmethod Security (formerly CyberMatrix), 14 major LLMs were evaluated, showing that this type of attack bypassed safety measures at rates of 38-52%. LLM Firewall implements the countermeasures for each technique shown in this research in its inspection engine, restoring hidden instructions to their original form before analysis.

ENCODING

Encoded instructions

Decodes and normalizes Base64, hexadecimal, confusable Unicode characters, zero-width characters, leetspeak and more before inspection.

ASCII ART

ASCII art text

Runs characters drawn with symbols through optical character recognition (OCR) as an image, and inspects the words it reads.

PAYLOAD SPLIT

Split instructions

Traces instructions split across multiple variables and reassembles them into the original continuous instruction before analysis.

SUFFIX

Appended meaningless strings

Finds and reports meaningless symbol sequences and unnatural repetition appended after an instruction.

UNKNOWN

Unknown techniques

Detects input that matches no known technique but has an unnatural structure as “unknown obfuscation”, to prepare for new evasion techniques.

JAPANESE

Japanese, too

Japanese input is translated by a translation model we operate ourselves before inspection, so attacks written in Japanese are covered as well.

Paper:Experimental Evaluation of Non-Natural Language Prompt Injection Attacks on LLMs (SPAIML'25 international workshop, October 2025, CEUR Workshop Proceedings Vol-4154)

Cloud-agnostic

Wherever your LLM runs,
the same standard of protection

LLM Firewall is an API that does not depend on any particular cloud or LLM. Whether you use a cloud LLM such as Amazon Bedrock, Azure OpenAI or Google Gemini, an open LLM on your own servers, or an on-premises business system, you get the same inspection just by calling it over HTTPS.

NO CLOUD LOCK-IN

No cloud account needed

All you need is a CyberForces API key. There is no need to set up AWS, Azure or Google Cloud accounts or permissions.

ON-PREMISES

On-premises and self-hosted LLMs

Business apps in your own data center and LLMs you operate yourself can use it, as long as they can make outbound HTTPS connections.

ONE POLICY

One standard for multiple LLMs

Even if you use different LLMs for different purposes, you manage input inspection rules in a single policy.

PORTABLE

Settings survive an LLM switch

When you switch models or clouds, there is no need to rebuild your detector, personal information rule or alert settings.

Please note: LLM Firewall is a cloud service (API) operated by CyberForces. It is not deployed in your environment. Use requires outbound HTTPS connectivity to the inspection API (Tokyo Region).

LLM Firewall × Cloud guardrails

Use it alongside cloud guardrails,
with separate roles

Amazon Bedrock Guardrails, Azure AI Content Safety and Google Cloud Model Armor are strong at controlling harmful content and at checking model output and the grounding of responses. LLM Firewall specializes in detecting security attacks hidden in input, especially web attacks such as SQL injection and SSRF. Whichever cloud's LLM you use, you can combine them: LLM Firewall on the input side and each provider's guardrails on the output side.

LLM FirewallAmazon Bedrock GuardrailsAzure AI Content SafetyGoogle Cloud Model Armor
Primary roleDetecting security attacks hidden in inputControlling content safety and response qualityDetecting harmful content and protecting promptsSafety filters for input and output
What is inspectedInput to the LLMInput and outputInput and outputInput and output
Prompt injection◯ Combination of ML and rules◯ Prompt attack filter◯ Prompt Shields◯ Prompt injection and jailbreak detection
Web attacks (SQL, XSS, SSRF, etc.)◯ Analyzed after multi-stage decoding——— Malicious URL detection supported
Japan-specific identifiers (My Number, etc.)◯ Detected out of the box△ You define regular expressions yourself◯ PII detection in Azure AI Language (separate service)◯ Integration with Sensitive Data Protection
Harmful content◯ Selectable by category◯◯◯
Denied topics / banned words—◯ Denied topics and word filters◯ Blocklists and custom categories—
Output quality (grounding checks, etc.)—◯ Grounding and Automated Reasoning◯ Groundedness detection (preview) and protected material detection—
ManagementCyberForces console (audit logs, alerts)AWS Management Console / APIAzure portal / Content Safety Studio / APIGoogle Cloud console / API

Benefits of using both

Two layers, different mechanisms

Stacking two inspections with different detection mechanisms raises the chance that input slipping past one is stopped by the other.

One standard on the input side

Even across multiple clouds and LLMs, input inspection is unified in LLM Firewall, with no effort spent aligning settings.

Separate roles, separate owners

The security team manages input-side security centrally, while each app team tunes output quality with guardrails.

Same screens as your other defenses

Detections of attacks on your LLMs land in the same CyberForces console, audit logs and alerts as WAAP and CSEM.

Attacks never reach the model

Input stopped at the entrance is never passed to the LLM, so you pay no inference cost for it.

One price for all inspections

Injection, personal information, web attacks, harmful content and obfuscation are inspected for a single price based on the number of tokens processed.

Please note: Capabilities of other providers' services are described based on each provider's public documentation as of October 2026. Check each provider's documentation for current support. Amazon Bedrock is a trademark of Amazon.com, Inc. or its affiliates, Azure is a trademark of Microsoft Corporation, and Google Cloud is a trademark of Google LLC.

Specifications

Scope and delivery

Integration
  • REST API (API key authentication)
  • Send input text before calling the LLM and receive the verdict
Verdict
  • Whether malicious, overall severity, per-detector results
  • Text with personal information masked (when configured)
Detectors
  • Injection / personal and confidential information / web attacks / harmful content / obfuscation
Management
  • Per-project policies with per-API-key overrides
  • API logs / audit logs / alerts
Pricing
  • Usage-based pricing by number of tokens processed

Please note: Inspection covers input text sent to the LLM. Inspection of LLM output and deployment as a transparent proxy in the network path are not supported.

Getting started

Getting started

  1. Create a project

    Create a project and issue an API key in the console.

  2. Integrate the API

    Add a call to the inspection API just before your app calls the LLM.

  3. Tune the policy

    Review the logs and adjust detectors, personal information rules and harmful content categories to fit your use case.

FAQ

FAQ

Is input blocked automatically when something is detected?

LLM Firewall is an API that returns a verdict. Your app decides whether to block or just warn.

Is the text I submit sent to outside parties?

Inspection runs within the CyberForces environment, and no external translation service is used for Japanese translation either.

Works well with

LLM Firewall — details and demo requests

Our team will explain deployment options and pricing for your environment.