Show HN: Shoehorn, a library to quantize an LLM to fit your Mac's VRAM

  • Posted 53 minutes ago by rhgraysonii
  • 6 points
https://github.com/notactuallytreyanastasio/shoehorn
I made this after seeing someone posit the idea online yesterday over lunch then spent some time refining it. So far it's pretty impressive IMO! Right now I am running Qwen3-30B-A3B on my 24gb unified memory m4 MacBook Pro at 50 tok/sec and this should definitely not be working for such a large model on my middling hardware.

Things are detailed in the README to get up and running and DESIGN.md has details on all the choices and such made along the way.

0 comments