./resources / blog

AI Security: Attacking and Defending LLM Systems

Large language models have moved from novelty to core infrastructure, and with them comes a new attack surface that traditional application security does not fully cover. The OWASP Top 10 for LLM applications gives us a shared vocabulary — prompt injection, data poisoning, model extraction and more. This is a practitioner's map of those risks and the controls that contain them.

When an application embeds a language model, it inherits a peculiar property: the same channel carries both data and instructions, and the model cannot reliably tell one from the other. That single fact is the root of most LLM security problems. The OWASP Top 10 for LLM Applications organises the resulting risks into a checklist that security teams can actually work through. Below we take the most consequential entries, explain the mechanism, and pair each with the defensive control that matters most.

LLM threat model & defensive ring Input Validation Output Filtering Least-Privilege Tools Monitoring Human-in-loop Prompt Injection Data Poisoning Model Extraction Sensitive Disclosure Insecure Output Handling LLM System Preemptive Cyber Security
Figure 1. The LLM system sits at the centre of an inward-pointing threat ring; the outer ring of controls — validation, output filtering, least-privilege tools, monitoring — is what keeps the threats from reaching it.

Prompt injection — the defining LLM risk

Prompt injection exploits the model's inability to separate trusted instructions from untrusted input. In the direct form, a user crafts input that overrides the application's system instructions. The more dangerous indirect form hides instructions inside content the model later ingests — a web page, an email, a document in a retrieval pipeline — so that merely processing attacker-controlled data hijacks the model's behaviour.

Defending against injection

There is no single filter that fully solves this, so defence is architectural. Treat all model input as untrusted; keep privileged instructions out of reach of user-influenced text; constrain what the model is allowed to do rather than just what it is told; and require human confirmation for consequential actions. The mindset shift is to stop trusting the model's output as if it came from your own code.

Insecure output handling

The mirror image of injection is trusting what the model produces. If model output flows unchecked into a browser, a shell, a database query, or another system, then classic web vulnerabilities — cross-site scripting, command or SQL injection — reappear, now driven by text an attacker may have influenced. The control is old and reliable: encode, escape, and validate model output at every boundary exactly as you would any other untrusted data. An LLM is not a trusted component.

Treat the model as a confused deputy: powerful, eager to help, and unable to tell your instructions from an attacker's. Every control follows from that assumption.

Training-data poisoning

Models learn from data, and if an adversary can influence that data — during pre-training, fine-tuning, or through a retrieval corpus — they can plant biases, backdoors, or degraded behaviour that surface later. Data poisoning is a supply-chain problem: the mitigations are provenance and integrity for training sources, vetting of third-party datasets and models, anomaly detection over training data, and clear separation between trusted and community-sourced content.

Sensitive-information disclosure

Models can leak. They may reproduce sensitive material memorised from training data, or reveal confidential context supplied earlier in a session or through retrieval. Sensitive-information disclosure is contained by minimising what the model is ever exposed to: scrub secrets from training and prompt context, apply least-privilege to any data the model can retrieve, and filter outputs for sensitive patterns before they reach a user. The safest secret is the one the model was never given.

Model extraction and abuse

A model exposed through an API is an asset that can be probed. Model extraction attempts to reconstruct a proprietary model's behaviour, or infer properties of its training data, through large volumes of carefully chosen queries. Related abuses include denial-of-wallet, where an attacker drives up inference costs. Rate limiting, query monitoring for extraction-like patterns, and authentication that ties usage to accountable identities are the practical defences.

Excessive agency

As LLMs are wired to tools — sending email, calling APIs, executing code — the blast radius of a successful injection grows. Excessive agency is the risk of granting a model more capability, autonomy, or permission than the use case requires. The remedy is squarely a security-engineering one: give each tool the narrowest possible scope, prefer read-only where you can, and insert human approval in front of any irreversible or high-impact action. Least privilege applies to models exactly as it does to service accounts.

Building a defensive posture

Read across these entries and a coherent posture appears. Validate everything going in; never trust what comes out; lock down the data and tools the model touches; and monitor the whole system for the abuse patterns above. None of these is exotic — they are the familiar disciplines of input validation, output encoding, least privilege, supply-chain integrity, and logging, applied to a component that happens to speak natural language. Organisations should threat-model their LLM features as deliberately as any other critical system, and test them with the same rigour, drawing on resources from OWASP and emerging AI risk frameworks from bodies such as NIST.

Key takeaways

  • The core flaw is that an LLM cannot separate instructions from data — every control follows from that.
  • Prompt injection (direct and indirect) is the defining risk; defend it architecturally, not with a single filter.
  • Never trust model output — encode and validate it at every downstream boundary.
  • Data poisoning and sensitive disclosure are supply-chain and data-minimisation problems.
  • Apply least privilege to model tools and use the OWASP LLM Top 10 as your checklist.
#ai-security #llm #owasp #ml
All articles Secure your AI systems