Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I have no idea how artificial analysis got to be something anyone took seriously. This is their new benchmark set?

A glance at their new index shows that whatever they're measuring, it isn't useful.

Spend an hour with gemini 3.8 and tell me that model belongs in 2026. It feels like the model has Alzheimer's. It gets confused about whether what it reads is what it did. Just crazy bad.

I haven't tried muse spark 1.3. But it must have been a miracle since 1.2 to hit that rank.

Video game journalism vibes all over this.

 help



What other benchmarks do you recommend that are more accurate?

Give a task you have to 3 different models and see what actually works for you. There are no good benchmarks.

As the saying goes, Artificial Analysis is the worst benchmarking company, except for all the others.

Because it's the best option currently available.

It's really easy to shit on AI benchmarks, but that noise is useless unless you're offering a solution or a better benchmark.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: