Apple Silicon build for large models
A quiet, low-power route to running large models at home, built around a Mac Studio with plenty of unified memory. Slower generation than a discrete GPU, but able to hold models that a single consumer card cannot.
Key hardware
- Mac Studio M2 Ultra (64GB) — 64GB
This build suits someone who wants to run large models at home without the noise, heat, and power draw of a multi-GPU rig, and who is willing to accept slower generation in exchange.
Who it is for
Unified memory means most of the machine’s RAM is available to the model, so a 64GB Mac Studio can hold models that a single 24GB card cannot. That is the whole case for it. If your priority is the fastest possible tokens per second on mid-sized models, a discrete GPU will serve you better for the money.
What to expect
- Model sizes: comfortably runs large models that need more than 24GB, at reasonable quantisations.
- Speed: generation is slower than a high-end NVIDIA card, and prompt processing on long contexts is noticeably slower still. This is the main trade-off, and it is a real one.
- Living with it: near-silent and low-power, which makes it pleasant to have running on a desk.
Memory is fixed at purchase and cannot be upgraded later, so decide up front how much headroom you want. Buying more memory than you need today is the only way to leave room for larger models tomorrow.
Build last reviewed 15 January 2026.