Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
LoganDark
5 days ago
|
parent
|
context
|
favorite
| on:
Qwen 3.8 27B available on Cerebras at 1500 tokens/...
GPUs can't reach these speeds. You could build a supercomputing cluster and still not reach these speeds.
help
gpugreg
5 days ago
|
next
[–]
MiMo-V2.5-Pro-UltraSpeed gets pretty close with over 1000 TPS on 8x B200. It has 1.02T total parameters and 42B active, compared to 27B total/active for Qwen3.8-27B. Also, B300 are out now. I think 1500 TPS for Qwen3.8-27B should be doable.
reply
LoganDark
5 days ago
|
parent
|
next
[–]
That model uses a
lot
of tricks to achieve 1000 t/s. I would not use raw parameter counts alone for such comparisons, in general.
reply
nostrebored
5 days ago
|
prev
[–]
Celeris reaches ~50% of the speed on commodity hardware. celeris-magnus-1 is based on qwen3.8-27b. Maybe we will get there without custom silicon!
reply
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: