๐Ÿ’คQuietscore 0.0Aug 17, 2026ยท2608.16834cs.CLcs.AI

Model Hypnosis: Strong control of AI via additive subliminal effects

Enric Boix-Adsera, Benedict Tessler

Narrative

No narrative written yet. The narrate cron picks top papers by score; run /api/cron/narrate to populate this manually.

Abstract

We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irrelevant cues in the prompt can be systematically combined to strongly control model behavior. Model hypnosis occurs across model families and scales, including in frontier reasoning models, and hypnotic prompts can transfer between models. Because the model is controlled by inconspicuous textual choices, such as paraphrases and typos, model hypnosis presents new challenges and avenues for AI safety, and is a major hurdle for AI interpretability.

Citation timeline
Not enough citation snapshots yet to plot a timeline. Come back after a few cron runs.