Frontier models went from not being able to count the number of 'r's in "strawberry" to getting gold at IMO in under 2 years [0], and people keep repeating the same clichés such as "LLMs can't reason" or "they're just next token predictors".
At this point, I think it can only be explained by ignorance, bad faith, or fear of becoming irrelevant.
> At this point, I think it can only be explained by ignorance, bad faith, or fear of becoming irrelevant.
Based on the past history with frontier-math & AIME 2025 [1],[2] I would not trust announcements which cant be independently verified. I am excited to try it out though.
Also, the performance of LLMs was not even bronze [3].
Finally, this article shows that LLMs were just mostly bluffing [4].
At this point, I think it can only be explained by ignorance, bad faith, or fear of becoming irrelevant.
[0] https://x.com/alexwei_/status/1946477742855532918