A 4,000-token context is enough for quick help, but it makes the laptop LLM a handy sidekick rather than a stand-in for a full research workflow.
I run an LLM offline on my laptop. A heavily quantized Qwen3.6, using LMStudio. MBP M4 Max 36GB RAM.
It does 8-12 tokens per second. The context window is quite small (like 4000 tokens), but it's good enough for quick science/tech questions or helping me do a high-level grammar and narrative check while writing.
I use it on flights with no wifi. It eats a LOT of battery life and gives me an eerie scifi feeling. I can tap all human knowledge, with no internet connection, in my "above average" laptop, compressed into 20GB.
It's as easy as installing LMStudio and downloading the weights from HuggingFace.