Retracted: public red-team reports
This post announced public red-team reports on capable models. The programme does not exist. Kept as a correction rather than deleted, with what actually runs.
By Safety & Trust Team

This post previously announced that every model above a capability threshold ships with a public red-team report on its listing page, covering a standard evaluation suite with links to the underlying traces.
That was not true, and none of it was built. There is no capability threshold, no evaluation suite, no red-team programme, and no safety report attached to any listing page. We are retracting the post rather than deleting it, because an announced safety control is precisely the kind of thing someone plans a deployment around, and a silent deletion would leave anyone who read it none the wiser.
What actually runs
- Format validation on every upload: the artifact must be structurally what it claims to be. This is a correctness check, not a security check.
- An automated policy screen on AI Factory-generated assets only, reading the prompt and generated artifacts for surveillance use, kinetic harm, non-consensual biometric identification, CSAM, malware, embedded secrets, and license conflicts. A block verdict does prevent publishing.
- A moderation queue and takedown path, operated by people and largely reactive.
What that means for you
Treat every artifact you download as untrusted code and sandbox it. If you need adversarial assurance for a model you are putting near people or hardware, you have to run that yourself — we are not doing it for you, and until now this site implied otherwise.
The safety-and-moderation guide is now written to be read literally, and lists what is not implemented alongside what is.


