Labs know the only thing the public even discusses on model releases are benchmarks, so they devote a majority of training on just benchmaxxing. It’s marketing.
Muse Spark looks great in benchmarks. Everyone I know who has tried it (Rust & C++ projects) has determined it’s a resounding “meh”. That doesn’t mean it’s not a great tool for frontend devs, I wouldn’t know.
They are literally frequently discussed and it's why there are different benchmarks for different domains.