AI Security Documentation
This section provides a structured reference for securing AI systems across the lifecycle: from model development through deployment and ongoing operations. Every page cites final, in-force instruments — NIST publications, ISO standards, EU regulations, OWASP and MITRE frameworks — with resolvable URLs and specific provisions.
The landscape moves quickly. This documentation is updated as standards evolve; the revision history for each page is available via Git.
Standards & Frameworks
The normative foundations for AI security and governance.
- NIST AI Risk Management Framework (AI RMF 1.0) — Core, Profiles, and Playbook structure
- NIST Generative AI Profile (AI 600-1) — GAI-specific risk mapping and controls
- ISO/IEC 42001:2023 — AI management system requirements and Annex A controls
- EU AI Act — Regulation (EU) 2024/1689: prohibited practices, high-risk obligations, GPAI rules, and phased enforcement
- OWASP LLM Top 10 (2025) — LLM01–LLM10 taxonomy with mitigation patterns
- MITRE ATLAS — Adversarial Threat Landscape for AI Systems: tactics, techniques, and case studies
- Standards Comparison — Mapping overlaps and gaps across AI RMF, ISO 42001, EU AI Act, and OWASP
Threat Taxonomy
Catalogue of attack vectors organised by MITRE ATLAS alignment where applicable.
- Prompt Injection — Direct, indirect, and stored injection; boundary confusion; data exfiltration paths
- Jailbreaks — Role-play, encoding, multi-turn, and many-shot techniques; defence evasion
- Adversarial Attacks — GCG, PAIR, TAP, and gradient-based methods against open and closed models
- Data Poisoning — Clean-label, backdoor, and availability poisoning; supply-chain vectors
- Model Extraction — Query-based extraction, distillation, and membership inference
- Supply-Chain Attacks — Compromised dependencies, model hubs, and CI/CD pipelines
- Agent Hijacking — Tool misuse, context window manipulation, and multi-agent delegation abuse
Defensive Architectures
Practical controls mapped to threat vectors and standards requirements.
- Input Guardrails — Prompt validation, instruction hierarchy, and classifier-based filtering
- Output Filtering — PII redaction, groundedness checks, and refusal enforcement
- RAG Security — Vector store access control, retrieval authorisation, and context injection defence
- Runtime Monitoring — Anomaly detection, drift alerts, and behavioural baselines
- Secure Deployment — TEEs, confidential computing, model signing, and supply-chain verification
- Defence in Depth — Layered control strategy aligning NIST AI RMF functions to technical controls
Governance & Operations
Operationalising standards into auditable processes.
- AI Risk Assessment — AI RMF Map function: context, stakeholder, and risk identification methodologies
- AI Incident Response — Detection, containment, eradication, and post-incident reporting per ISO 42001 Annex A
- Audit & Evidence — Artefact collection, model cards, data sheets, and conformity assessment preparation
- Transparency & Documentation — System cards, dataset documentation, and EU AI Act Art. 50/53 disclosure obligations
Engagement
- Request an AI Security Assessment — Scope, methodology, delivery via NestVault365, and contact