Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I wouldn't exactly call it snappy, but faster than Pro, yes.


Single request depth on vllm with dspark, I'm getting ~200 tps, I'd say it's pretty snappy.


Well sure but you're running on tens of thousands of dollars of hardware.


It's much faster than other models on that same hardware in the same size class. I've tested a few, it's by far the fastest I've tested.

And it wasn't tens* until recently. Didn't expect this to be one of my best performing assets this year.


I get like 80.


What's your setup? Happy to try to point you in the direction that worked for me.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: