AI Security Consulting
Threat modeling for LLMs, data pipelines, and model endpoints — plus the controls to close the gaps.
AI systems fail in ways the existing security programme was not built to catch. A model endpoint is an API that accepts natural language and returns actions, which makes the classic separation between data and instruction difficult to enforce. An agent with a credential is a principal your IAM model probably does not represent.
The result is a gap that neither the application security team nor the cloud security team fully owns, and it usually surfaces during a customer security review rather than internally.
This engagement threat-models the AI system specifically and returns the controls that close what it finds, in severity order.
What the engagement includes
AI-specific threat model
Prompt injection, both direct and indirect through retrieved content; tool and function-call abuse; training and retrieval data poisoning; model extraction; and excessive agency.
Agent identity and authorisation review
Whether each agent is a distinct principal with scoped credentials, and whether its actions are authorised individually or inherited from a broad service account.
Data pipeline and endpoint review
Where sensitive data enters the system, where it is retained, what a model provider receives, and whether tenant isolation holds under a crafted request.
Control recommendations
Specific, implementable controls ranked by severity — policy-as-code authorisation, retrieval isolation, output handling, and detection you can operate.
Framework mapping
Findings mapped to NIST AI RMF, ISO 42001, and your existing control framework, so the work counts toward obligations you already carry.
Who this is for
- Teams shipping LLM or agent features into production
- Security functions asked to sign off on an AI system for the first time
- Organisations facing customer security reviews that now ask about AI
When this is the wrong engagement: This is an application and platform security engagement. It is not a penetration test of your wider estate, and it does not replace one.
How it runs
- Discovery (30 minutes). A working call to map your stack, constraints, and the highest-value first step. No pitch deck.
- Scoped fixed-price proposal. A written statement of work with deliverables, milestones, and a fixed price — not an open-ended time-and-materials meter.
- Delivery. Senior architects execute against milestones, with governance applied to every AI-agent action from day one.
- Operate or hand off. Monitored operation, or a clean handoff with runbooks so your team can run it without us.
Questions we get asked
Is Citadel certified in the frameworks you assess against?
No, and the distinction matters enough that it is published on our security page. Citadel delivers against SOC 2, ISO 27001, HIPAA, FedRAMP-aligned, and CMMC control expectations, and maps findings to NIST AI RMF and ISO 42001. We do not represent ourselves as holding those certifications. Any consultancy that blurs that line is telling you something about how it treats other claims.
Can prompt injection be fixed?
Not eliminated, and treat anyone who says otherwise with caution. It is mitigated architecturally: assume model output is untrusted, authorise every action against policy rather than trusting the model to have decided correctly, and isolate retrieval so a successful injection cannot reach data the user could not already access. The realistic objective is bounding the blast radius, not preventing every injection.
Do you test against our live system?
Only with written authorisation and an agreed scope and window. Most of the value is in design review and threat modeling, which requires no live testing. Where hands-on validation is in scope, it is a defined exercise against defined targets, and the rules of engagement are agreed before anything runs.
What do we get at the end?
A threat model, findings ranked by severity with concrete remediation for each, a framework mapping, and a readout session with your security and engineering teams together. The deliverable is written for both audiences, because AI security findings that only the security team understands do not get fixed.
Related services
AI Governance
Identity per agent, autonomy tiers, and audit trails your risk and security teams will actually accept.
Enterprise AI Architecture
Reference architectures for agents, RAG, and MLOps that hold up under real load, security review, and audit.
Multi-Agent Systems
Fleets of governed agents coordinated around real workflows — not a demo that stalls after the meeting.
Start with a 30-minute discovery call
No pitch deck. We map your constraints and tell you the highest-value first step — including when that step is not an engagement with us.
Book a discovery call