llm-hasher: Local PII Detection and Tokenization for Hybrid LLM Workflows

✍️ OpenClawRadar📅 Published: April 15, 2026🔗 Source
llm-hasher: Local PII Detection and Tokenization for Hybrid LLM Workflows
Ad

llm-hasher addresses a specific security gap in hybrid LLM workflows: when you run local LLMs but still call external services like OpenAI, Claude, or Gemini for certain tasks, your PII still leaves your infrastructure in plaintext. This tool runs PII detection entirely locally using Ollama, so no data leaves your systems during the detection phase.

How It Works

The process follows three steps: detect PII locally, tokenize it before external LLM calls, then restore the original values after processing. This prevents sensitive data from being exposed to third-party services.

Detection Approach

The detection system uses a hybrid approach:

  • Regex patterns for structured data types: credit cards, IBAN numbers, email addresses, and IPv4 addresses
  • Ollama with llama3.2:3b (by default) for contextual detection of unstructured PII: names, addresses, national IDs, passports, and dates of birth

Technical Implementation

Mappings between original PII and tokens are stored in an AES-256-GCM encrypted SQLite vault. Deployment is simplified with Docker Compose, which spins up both Ollama and the llm-hasher service with a single command.

📖 Read the full source: r/LocalLLaMA

Ad

👀 See Also

ClawGuard: Open-Source Security Gateway for OpenClaw API Credential Protection
Security

ClawGuard: Open-Source Security Gateway for OpenClaw API Credential Protection

ClawGuard is a security gateway that sits between AI agents and external APIs, using dummy credentials on the agent machine while storing real tokens separately. It provides Telegram approval for sensitive calls and maintains an audit trail of requests.

OpenClawRadar
AI Agents Enable Solo Hackers to Breach Governments and Ransomware Campaigns
Security

AI Agents Enable Solo Hackers to Breach Governments and Ransomware Campaigns

A solo operator using Claude Code and ChatGPT exfiltrated 150 GB from Mexican government agencies, including 195 million taxpayer records. Another attacker used Claude Code to run an end-to-end extortion campaign against 17 healthcare and emergency services organizations.

OpenClawRadar
Claude Code Finds 23-Year-Old Linux Kernel Vulnerability
Security

Claude Code Finds 23-Year-Old Linux Kernel Vulnerability

Anthropic researcher Nicholas Carlini used Claude Code to discover multiple remotely exploitable heap buffer overflows in the Linux kernel, including one that had been hidden for 23 years. The AI found the bugs with minimal oversight by scanning the entire kernel source tree.

OpenClawRadar
Zero-Trust OpenClaw Architecture Adds Pre-Execution Authorization and Post-Execution Verification
Security

Zero-Trust OpenClaw Architecture Adds Pre-Execution Authorization and Post-Execution Verification

An open-source architecture for OpenClaw adds two security checkpoints: a Rust sidecar that intercepts tool calls before execution with sub-millisecond authorization overhead, and deterministic post-execution verification using assertions instead of LLM judgment. The system includes tracing with DOM snapshots and screenshots, plus a DOM compression skill that reduces token usage by 90-99%.

OpenClawRadar