But I'm also wondering about being able to run them on consumer-end hardware.
I remember using a local model 2-3 years ago and had to wait around 2-3 minutes for a basic answer to be printed. Now I'm running a "thinking" Qwen on a 16GB GPU and I'm able to do anything I'd do with Opus a couple months ago, at nearly the same speed. But that does use my entire VRAM and most of the RAM I have. No way I can also run a game or something else on the side.
But like how we went from bulky PCs to smartphones 1000x faster, and at the rate local models already improved, do you think we'll ever be able to have the same kind of models running locally, on affordable hardware, on our phones or maybe our fridges?
Not saying we should use them on anything, that will be a question for later, but strictly thinking about capabilities.