Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's interesting that all three of those used roughly the same amount of tokens, and almost entirely output. Feels like the thinking level lever didn't alter cost at all for this specific task, even though it did change the output.


I never trust OpenRouter to forward parameters correctly and would only ever conduct benchmarks with the official api, personally.


should use an open source one you can audit like my site TrustedRouter


That raises the question of what is it actually doing?

If it isn't spending tokens on quality, is it the assumptions about the task difficulty that cause it to perform better? Or are their broader differences in the model being run.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: