Trefon.
AI & Automation

Security Blocked Our AI Project. Now What?

Your AI pilot worked. Security said no to sending data to the cloud LLM - and they're right. Here are the options and how to start that conversation.

Jacek Trefon · · 8 min

TL;DR

Three deployment models for security-conscious AI: self-hosted open-source (Llama, Mistral) for maximum data control, compliant cloud (Azure OpenAI, AWS Bedrock) with data residency and no-training clauses, or hybrid minimizing data exposure. Start the conversation with security before building - retrofitting compliance after deployment is 3x harder.

The blocker nobody predicted

Your AI pilot worked. The demo was impressive. Everyone was excited.

Then security got involved. “We can’t send customer data to an external API.” Or “Our compliance team needs to review this.” Or “The data protection officer has concerns about GDPR.”

Your project stalls. The model works, but the path to production is blocked.

I see this pattern in nearly every AI project I audit. And here’s the hard truth: your security team is usually right. Sending customer data to a third-party AI API without due diligence is a real risk. The question isn’t “how do I bypass security?” - it’s “how do I build for compliance from the start?”

Three deployment models

Model 1: Self-hosted open-source (maximum data control)

Run an open-source model on your own infrastructure. No data leaves your network. Llama, Mistral, and other models can be deployed on your servers, in your cloud account, or on-prem.

Best for: Regulated industries (finance, healthcare, legal), companies with strict data residency requirements, use cases involving PII or sensitive data.

Costs: Higher upfront: €10-30K for infrastructure setup, €1-5K/month for compute. No per-API-token costs, but significant operational overhead.

Trade-offs: Lower accuracy than top-tier commercial models for most tasks. You need ML engineering talent to set up and maintain. The model won’t improve automatically - you manage updates yourself.

Model 2: Compliant cloud deployment (balanced)

Use a commercial model through a cloud provider that offers data processing agreements that protect your data. Azure OpenAI, AWS Bedrock, and GCP Vertex AI all offer options where your data is not used for training and stays within your chosen region.

Best for: Companies that want high accuracy without sending data to arbitrary third parties, need EU data residency, are already on a major cloud provider.

Costs: €5-15K setup, €200-2K/month in API fees plus cloud infrastructure. Higher accuracy than open-source for most tasks.

Trade-offs: You’re still sending data to a cloud provider. While contractual protections exist, data technically leaves your infrastructure. Some industries (defense, certain financial services) can’t accept this.

Model 3: Minimized data exposure (compromise)

Redesign the system so that sensitive data never reaches the AI. Use anonymization, pseudonymization, or local processing to strip PII before sending. Or use a hybrid model where sensitive data is processed locally and non-sensitive data uses cloud AI.

Best for: Organizations where some data is sensitive and some isn’t, teams who want cloud AI accuracy but need to protect certain data categories.

Costs: €15-30K additional development to build anonymization layers, plus ongoing maintenance of the data filtering pipeline.

Trade-offs: Adds architectural complexity. You need to ensure the anonymization is thorough - partial PII exposure is worse than no protection.

The conversation you should have before building

Before you write a line of AI code, have this conversation with your security and compliance teams. It will save you months of rework.

Ask them:

  • What data categories are off-limits for external processing?
  • Do we need EU data residency?
  • Can we use cloud AI providers with contractual protections?
  • Is self-hosting required, or is a compliant cloud deployment acceptable?
  • What’s the approval process for a new AI system?

Then design for their answer.

If they say “no external AI APIs at all,” use self-hosted models. If they say “approved providers only,” pick a compliant cloud deployment. If they say “it depends on the data,” build a data classification layer.

The worst approach is building the AI system and then asking security for approval. By then, you’ve committed to an architecture that may not be acceptable.

The cost of non-compliance

Building without security approval and finding out later that your architecture is non-compliant is expensive:

  • Re-architecture: 2-6 months of work to switch from cloud API to self-hosted
  • Data migration: Moving training data, embeddings, and pipelines to new infrastructure
  • Re-certification: New security review, new compliance documentation, new approval process

Total cost: easily 3x the original build cost, plus 3-6 months of delay.

Building for compliance from the start costs 10-15% more upfront. Retrofitting costs 3x after the fact.

The air-gapped AI option

For the most regulated environments, consider fully air-gapped AI - an AI system that runs on infrastructure with no internet connectivity.

This means self-hosting the model, running inference locally, and never connecting to any external service. It’s the most secure option and the most operationally demanding.

Real examples:

  • A defense contractor running Llama on an on-prem server for document classification
  • A healthcare provider running Mistral in their own cloud account for clinical note processing
  • A financial services firm running a fine-tuned model on air-gapped infrastructure for fraud detection

If this is your requirement, budget for operational overhead. Air-gapped AI isn’t a “set and forget” system - it requires ongoing maintenance, monitoring, and updates that you manage yourself.

FAQ

Can we use ChatGPT and still be compliant? For public data, yes. For customer data, probably not - your data is sent to OpenAI and may be used for training unless you have a business agreement that opts you out. Check your data processing agreement before sending any customer data to ChatGPT or the OpenAI API.

What about GDPR? Under GDPR, you’re responsible for data your AI system processes, even if it’s processed by a third-party API. You need a data processing agreement with the AI provider, and you must ensure data residency requirements are met. A compliant cloud deployment (Azure OpenAI, AWS Bedrock) typically provides these protections. Direct API access (OpenAI, Anthropic) may not.

Does the EU AI Act change anything? The EU AI Act regulates how you use AI, not where you host it. The deployment model affects privacy and security (GDPR), but the Act’s requirements (risk classification, documentation, human oversight) apply regardless of whether you self-host or use a cloud API.

Ready to apply this to your situation?

Book an AI Readiness Call

30-min call. No pitch. You leave with one concrete next step - even if it’s not us.

Jacek Trefon

Jacek Trefon

AI engineering leader. 28 years building technology, 4+ years building production AI systems. I help companies assess, architect, build, and deploy AI that actually ships. Based in Spain, working globally.