It's interesting that all three of those used roughly the same amount of tokens, and almost entirely output. Feels like the thinking level lever didn't alter cost at all for this specific task, even though it did change the output.
That raises the question of what is it actually doing?
If it isn't spending tokens on quality, is it the assumptions about the task difficulty that cause it to perform better? Or are their broader differences in the model being run.
If I live my life on the basis that some people don't share my sense of humor, and hence I should avoid doing anything funny that might be misunderstood, my life will be a lot less fun.
You're doing great. Don't let insanely low-effort (negative-effort, as in making others dumber rather than having no effect?) comments like from the above throwaway affect your actions.
Wow, the low, medium, and high pelicans came out in surprisingly different styles: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...