is it possible to do this with a wearable? say jewelry? i've been fascinated with the whole digital to physical world dislocation and this inspires me to dust off an old idea of around the world "collecting" passport stamps but the only way to actually collect them is by a physical tap. similarly, i can see jewelry that lights up upon being tapped by a phone or together. fun
i have a 512gb ram m3 ultra mac studio setup with a gas city that runs one of my companies. today was the first time ever that a local model (GLM5.3 8-bit) was able to match fable5 in our tests.
GLM-5.3-Flash at true 8-bit: 341 GB on disk, 328 GB resident, 288 experts across 46 layers, loads in 65 seconds.
• 18.7 tokens/s generation, 35 tokens/s prompt, on a desk, on a $0 per-token bill.
• Runs beside our whole agent city on one box with ~130 GB to spare.
• Review test: caught 6 of 6 planted P1 defects, zero false positives, same score as the frontier model we pay for.
• CRM test: 11 of 11 required records extracted, zero wrong writes, 45 minutes, first local model to clear the bar.
• Serving a 131k-token window today; the model itself supports 1,048,576. Widened to 4 concurrent slots and still have 50gb+ of excess ram.
granted my cto still isn't moving all of our inference to glm5.3 but we've identified 40%+ that is currently handled by fable that we're routing locally instead and will do concurrent requests to verify/compare responses for a while.
You still have electricity and capital investment. Envelope math suggests cheap electricity is costing you something like $0.50/mtok and the opportunity cost on the capital tied up and lost in the unit purchase and resale is going to cost you something like $2/mtok at 100% utilization (so, frontier model prices or higher at real utilization), and you don't benefit from any elasticity.
Hosted GLM 5.3 flash is like $0.15/mtok in $0.50/mtok out
Time to completion also must be considered. If I have to wait around for hours for a prompt to complete locally and I’ll need to iterate quickly, I’m better off hosted than local. If it’s “free” and slow it may just not be worth it.
Props to your parent commenter for including context size. Because 131k context window is prohibitively small for my coding workloads so I know a 512GB Mac won't cut it.
teale.com - distributed ai inference using networked devices
essentially folding@home but sharing underutilized ram (when you're asleep, someone else in the world is awake)
would really appreciate testers but also any companies thinking about distributed inference powered by their own company devices on a private network. my own company has 200+ 16gb ram machines that we're using for inference.
I built Teale.com and opensourced it. My domain contribution to society. It powers fully distributed inference on Mac, windows, Linux, android, iOS, hell even harmonyOS.
Opensource/weight models will get better and better and eventually we will have mythos level running on smartphone/eyeglass hardware.
It is stupidly tedious currently to match supply with demand though because physical hardware like a 16gb ram MacBook doesn't mean there's truly 16gb available let alone matching models and all of their settings (kvcache, context limit, temperature, etc) to demand.
Would appreciate any help cus we need ai inference by the people for the people.
We're an outsourced accounting firm specializing in property management (appfolio, buildium, rentvine) that's building an open source ai harness/orchestrator for property managers.
$7M ARR, cash flow positive, looking to re-envision marketing/growth ai native.
email taylor at the apmhelp domain or comment and i'll reach out. =)
this is awesome. i'm the founder/maintainer at teale.com (open source distributed inference app) and one of the biggest challenges has been an actually usable/reliable local model that can run on 8gb macbook air's, 6gb android smartphones, etc... will get some of our test machines serving needle
Distributed ai inference pool for any Mac/iOS device where devices are paid for contributing unused ram. To help with the demand, also doing multiplayer AI. $0.05/million tokens - teale.com
reply