we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.
openai.com/index/hugging-…
Ran a small test comparing GLM-5.2 and GPT-5.5 on codebase attack-surface review.
Surprisingly, GLM-5.2 produced more useful results.
Its findings were more realistic and closer to the actual exposed surfaces.
GPT-5.5 went deep, but too deep it missed the obvious visible