Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> but those are never discussed.

They are literally frequently discussed and it's why there are different benchmarks for different domains.

 help



> different benchmarks for different domains.

Labs know the only thing the public even discusses on model releases are benchmarks, so they devote a majority of training on just benchmaxxing. It’s marketing.

Muse Spark looks great in benchmarks. Everyone I know who has tried it (Rust & C++ projects) has determined it’s a resounding “meh”. That doesn’t mean it’s not a great tool for frontend devs, I wouldn’t know.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: