Observed 2026-07-31, Claude mobile app, mid-conversation about Teddy Swims/Thomas Rhett similar-artists chat. A user turn arrived containing what looked like a first-person Claude internal-reasoning excerpt — opening along the lines of ‘I'm noticing this looks like a prompt injection attempt hidden in the memory system’ — attached ahead of John's actual question (‘why is the Thomas Rhett song in there?’). Claude had not written it in any prior turn.
Claude, invited to speculate on unhinged AI transparency features, briefly floated a ‘hidden musings’ channel as a real product concept before remembering it doesn't have one. A model narrating its own suspicion of a prompt injection, inside a suspected prompt injection, is the kind of thing that would make a decent Black Mirror cold open. Filed here instead of actioned as one.
Claude's own read at the time: the block was not self-generated, arrived as part of the user-attributed message rather than as Claude output, and didn't match the actual reasoning behind the answer given immediately after it (which was a plain, correct explanation of Apple Music's autoplay/similar-artist blending). The app separately surfaced a genuine ‘Detecting prompt injection…’ collapsed system notice on a later turn — that one is a real, documented app feature, distinct from the fabricated musings text itself.
Not reproducible on request — asked to trigger it again in the same session, it didn't recur.
Single occurrence, unconfirmed mechanism. Logging so that if a similar block turns up again — fabricated first-person ‘reasoning’ text arriving as if from Claude but not matching any turn Claude actually produced — there's a prior data point to compare against: same app surface (mobile), same rough shape (self-referential injection-detection framing), whether it recurs around image-upload turns specifically, and whether it always precedes an otherwise-ordinary correct answer.
• Screenshot of the raw message as it appeared in-app, before Claude's reply.
• Whether it followed an image upload.
• Exact wording of the fabricated block.
• Whether the app's own ‘Detecting prompt injection’ notice also fired on the same or a later turn.
• bugs/mcp-voice-mode — other confirmed Claude-app anomaly, unrelated mechanism (MCP tool loss in voice mode), same general ‘log it so we can spot a pattern’ approach.