Updated
Updated · The Verge · Oct 11
Author Tests 125B-Parameter Qwen on $12,000 Mac Studio for Private Local AI
Updated
Updated · The Verge · Oct 11

Author Tests 125B-Parameter Qwen on $12,000 Mac Studio for Private Local AI

2 articles · Updated · The Verge · Oct 11

Summary

  • A 125-billion-parameter Qwen 3.8 Flash Next model, about 105GB in size, was installed through the self-hosted Hermes Agent on an M5 Ultra Mac Studio with 256GB of unified memory.
  • Privacy drove the setup: running models locally let the author analyze sensitive financial records and embargoed laptop specs without sending data to OpenAI, Google, Microsoft, Anthropic or other cloud services.
  • Hermes handled a few early tasks, including a 7:30 a.m. email-calendar-weather briefing and reorganizing a 400-plus-game Steam library by genre after temporary access via a Steam web API key.
  • The experiment also exposed limits: the daily briefing initially failed because macOS was asleep, later broke again, and a larger effort to automate repeated laptop benchmark tests remains unfinished.
  • The broader test is whether expensive high-RAM machines such as Apple’s AI-focused desktops can make local agents useful enough to justify their cost for privacy-conscious users.

Insights

As local AI takes over desktop chores, will the massive hardware costs outweigh the benefits of ditching cloud subscriptions?
Can running a massive AI locally truly keep your private data safe, or is it just an expensive illusion of security?
If an offline AI needs constant babysitting to organize files, is the dream of a private digital assistant already dead?