GPT-5.6 had no problem seeing the text when giving it a video recording of it.
"No problem" as in using temporal analysis with optical flow and vertical-displacement maps to estimate how the image moved, and combine those into a motion map with increased contrast to see the text. I didn't give it any instructions though, just asked it what it said.
For people thinking it detected the decoy text. It did not. It was the text I entered myself. Here are the thinking steps:
Inspected video metadata, extracted frames, and analyzed content
Detected and analyzed periodicity in stereogram image data
Blurred, downsampled, and analyzed image frames
Computed frame pixel shifts and cross-correlation
Checked OpenCV installation, optimized and analyzed optical flow processing
Analyzed motion for text
Estimated vertical displacement and processed image flow
Refined text extraction
Processed optical flow and analyzed image
I threw it at Fable with the prompt "What does the message say?" and it also figured out without issue. Thought process is interesting https://imgur.com/a/GToXs6W
That's because the "humans can read" part is not visible to all humans. Like myself I thought the image I could see was the one the machine couldn't. I managed to extract both forms of text by extremely crude means.
What this is, is a message that is more reliably readable by a sufficiently good AI than it is by all humans.
The chat contains the video, which may reveal som meta data about me. But the shared chat also doesn't include any of the images in the thinking steps that my original chat has, where you can see the different maps it generated. Pretty cool seeing it go from noise to patterns to the final image which clearly shows the text as black on white.
"No problem" as in using temporal analysis with optical flow and vertical-displacement maps to estimate how the image moved, and combine those into a motion map with increased contrast to see the text. I didn't give it any instructions though, just asked it what it said.