Multi-layer Data Anonymization Before Public LLMs
Client: Qlos
- Challenge
- Diagnostic data collected from Qlos's and its clients' systems (tickets, logs, documentation) carries personal data (names, e-mail addresses, phone numbers, national ID numbers, postal addresses, usernames), company identifiers (client and provider names, Polish tax and company registry numbers) and infrastructure and operational details (hostnames, IP addresses and ranges, ticket numbers). None of it could be sent to a public LLM as-is.
- What we built
- An anonymization pipeline with multiple independent detection layers, running inside the environment before any request reaches an external model. Beyond standard personal data, it covers IT-specific identifiers that generic PII tools miss.
- Result
- Public LLMs can be used on real operational data without exposing personal data or the infrastructure details of Qlos's clients.