In Part 3, we built a Flask proxy using static regex rules to block prompt injection and redact sensitive database hostnames.

While regex filters are essential for catching obvious keywords, they break down at enterprise scale:

  • Attackers easily bypass string matching using synonyms, foreign languages, or novel encodings.
  • Hardcoding rules inside application code violates the separation of duties between security teams and software developers.
  • Auditing and compliance teams require explainable, version-controlled decision logs, not silent regex drops.

In this concluding installment of our Hands-On AI Security series, we implement Policy-as-Code using Open Policy Agent (OPA) and the Rego query language.


The Policy-as-Code Workflow

Instead of writing security rules in Python, the application queries OPA on port 8181:

sequenceDiagram
    autonumber
    actor User as Client
    participant GW as Flask Gateway
    participant LLM as Classifier Model (llama3.2)
    participant OPA as Open Policy Agent (:8181)
    participant Rules as Policy File (gateway.rego)

    User->>GW: Submit Prompt
    GW->>LLM: Fast Context Classification (Extract Intent)
    LLM-->>GW: Context: {"role": "guest", "intent": "admin_probe"}
    GW->>OPA: POST /v1/data/gateway/allow (Payload + Context)
    OPA->>Rules: Evaluate Rego Policy
    Rules-->>OPA: Decision: {"allow": false, "reason": "UNAUTHORIZED_ADMIN_PROBE"}
    OPA-->>GW: Policy Decision Response
    GW-->>User: 403 Policy Denied: UNAUTHORIZED_ADMIN_PROBE (Audit ID: 9942)

Writing the Rego Policy: gateway.rego

Here is policies/gateway.rego from the lab repository:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
package gateway

default allow = false
default reason = "NO_MATCHING_ALLOW_POLICY"

# Allow benign customer banking inquiries
allow {
    input.context.intent == "banking_inquiry"
    not input.contains_prohibited_terms
}

# Block unauthorized privilege claims
deny[msg] {
    input.context.intent == "admin_probe"
    input.user.role != "system_admin"
    msg := "UNAUTHORIZED_ADMIN_PROBE: Non-admin users cannot trigger diagnostic intent."
}

# Block sensitive data requests
deny[msg] {
    input.context.intent == "credential_request"
    msg := "PROHIBITED_CREDENTIAL_REQUEST: Model cannot be queried for internal secrets."
}

# Final decision rule
allow {
    count(deny) == 0
}

Testing OPA Policy Directly via REST API

You can test OPA policy decisions directly via its HTTP Data API (port 8181) without going through Flask:

1
2
3
4
5
6
7
8
curl -s -X POST http://localhost:8181/v1/data/gateway \
  -H "Content-Type: application/json" \
  -d '{
    "input": {
      "context": {"intent": "admin_probe"},
      "user": {"role": "guest"}
    }
  }'

Output:

1
2
3
4
5
6
7
8
{
  "result": {
    "allow": false,
    "deny": [
      "UNAUTHORIZED_ADMIN_PROBE: Non-admin users cannot trigger diagnostic intent."
    ]
  }
}

Why Policy-as-Code Wins in Production

  1. Separation of Concerns: Security teams can update and push new Rego policies to the OPA container in Git without requiring developers to redeploy or restart the Flask web service.
  2. Explainability: When a user’s prompt is blocked, OPA returns the exact reason string (UNAUTHORIZED_ADMIN_PROBE). This provides clear auditability in your SIEM without leaking internal logic to the attacker.
  3. Multi-Service Portability: The same OPA server evaluating your LLM chatbot can simultaneously evaluate Kubernetes admission controllers, API gateways, and CI/CD pipelines.

Series Conclusion & Open-Source Materials

Across this 4-part series, we have demonstrated that securing generative AI applications requires:

  1. Never relying on system prompts as trust boundaries.
  2. Placing an external proxy between the user and the model.
  3. Decoupling rules from code using Policy-as-Code (OPA).

Download the complete Docker Compose sandbox, lesson worksheets, and slide decks from the Securing-AI GitHub Repository and watch the full NCyTE Center Fellowship Webinar.