I don't think a locally hosted LLM would be powerful enough for the supposed "ag...

lxgr · 2025-12-17T00:04:11 1765929851

Not yet, but we’ll hopefully get there within at most a few years.

Dylan16807 · 2025-12-17T02:50:43 1765939843

Get there by what mechanism? In the near term a good model pretty much requires a GPU, and it needs a lot of VRAM on that GPU. And the current state of the art of quantization has already gotten us most of the RAM-savings it possibly could.

And it doesn't look like the average computer with steam installed is going to get above 8GB VRAM for a long time, let alone the average computer in general. Even focusing on new computers it doesn't look that promising.

SirHumphrey · 2025-12-17T07:33:42 1765956822

By M series and amd strix halo. You don't actually need a gpu, if the manufacturer knows that the use case will be running transformer models a more specialized NPU coupled with higher memory bandwidth of on the package RAM.

This will not result in locally running SOTA sized models, but it could result in a percentage of people running 100B - 200B models, which are large enough to do some useful things.

Dylan16807 · 2025-12-17T07:49:39 1765957779

Those also contain powerful GPUs. Maybe I oversimplified but I considered them.

More importantly, it costs a lot of money to get that high bus width before you even add the memory. There is no way things like M pro and strix halo take over the mainstream in the next few years.

koolala · 2025-12-17T02:56:05 1765940165

This is probably their plan to monetize this. They will partner with a AI company to 'enhance' the browser with a paid cloud model and the local model has no monetary incentive not to suck.