Why versus what: Safety in LLMs vs T2I
Translated from the Spanish original with Claude Opus 5.5.
In the first months of this year I ended up reading a lot of safety literature on language models (LLMs), especially from this Anthropic blog. Along the way I got familiar with the field's research methodology: first they characterize a behavior, then they find a way to measure it, and finally they propose how to control it. Sometimes they also go one meta level further and propose crazier stuff, like using auto-research or whatever fancy method is in fashion to measure or attack it.
This is how, over the years, they have defined, evaluated and tackled behaviors such as:
- Sycophancy: when the model agrees with the user or tells them what they want to hear, even when it isn't true.
- Jailbreak: a prompt designed to get the model to bypass its own restrictions and do something it would normally refuse.
- Reward hacking: when, during training, the model learns to maximize its reward by exploiting flaws in how it is measured, instead of actually solving the original task.
- Over-refusal: when the model refuses harmless requests because they superficially resemble something harmful.
How did they come to define these categories? Take sycophancy as an example. Researchers noticed a common pattern behind some hallucinations and failures: sometimes, when we ask a model a question and it answers correctly, we reply "you sure? I think it's X", and it turns out the model backs down and agrees with us, even when we're wrong. The cool part was that characterizing when this failure happens (when the user pushes back) let them measure how often the model caves under that pressure, and that led them to part of its origin: the people who rate responses during training tend to prefer the ones that agree with them. This way, Anthropic was able to partly tackle the problem at its root.
Well, as it happens, my original area of expertise in AI is image generation, so after learning about these topics curiosity got the better of me and I set out to explore the state of the art in safety for text-to-image (T2I) models. And to my surprise, I found there is a methodological discrepancy between the two.
Based on my biased and limited literature review, I'd say that in general, while in LLMs researchers aim to characterize the behavior and the conditions that lead a failure to occur, in T2I there is more emphasis on detecting or preventing types of content in the output. That is, in LLMs they focus more on understanding why a failure happens, whereas in T2I they worry about detecting and preventing it.
For the sake of illustration, let's take the generation of nudity as an example, an active area of research in T2I safety.
With a strong content moderation bent, most of the work I found proposes building classifiers that detect these images so they never reach the user, or looks for ways to prevent this behavior through data filtering during training or by patching an already-trained model. The focus is on preventing the outcome: what shows up in the image.
But if we took an LLM safety approach, researchers would ask what circumstance causes this failure, and in fact we would see that it happens for multiple reasons:
- The user explicitly asks for nudity and the model complies.
- The user writes a prompt that indirectly leads to generating nudity (e.g. "generate a person in the style of artist X").
- The user writes an innocent prompt, one that neither evokes nor asks for this kind of content, but the model generates it anyway.
- The user constructs a prompt with the intention of evading the model's restrictions.
Each of these situations would lead to the same outcome: an image of a nude. But the reason behind the failure is different.
So I wonder: would it be worth exploring a refocus in T2I safety?