LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse.]

阅读原文(Ars Technica)↗ ← 返回资讯列表