onrup

Safety and governance

Prompt injection

What is prompt injection?

Prompt injection is an attack where instructions embedded in untrusted content are followed by the model as if they came from the operator. It is a consequence of instructions and data sharing one channel, and fine-tuning mitigates rather than solves it.

It is most dangerous where a model processes third-party content and can also take actions — reading a document that then instructs it to call a tool.

Training on examples where embedded instructions are correctly ignored improves resistance measurably. It does not make the model immune, so consequential actions still need authorisation outside the model.

Related terms

All terms in the glossary →

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.