Skip to content
local-ai

DeepSeek-R1-Distill-Qwen 32B

DeepSeek · Text generation · 32B · 131k context · Released 22 January 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

DeepSeek released a family of distilled models that transfer the reasoning behaviour of the full 671B R1 into smaller dense models. The 32B Qwen2.5-based version is the strongest of these that still runs on a single high-end card, and it inherits the base model's permissive Apache 2.0 licence.

Strengths

  • Strong reasoning for a model that runs on a single 24GB card
  • Inherits Qwen2.5-32B's Apache 2.0 licence
  • A practical route to R1-style reasoning without a server

Weaknesses

  • A distill, so it does not match the full 671B R1
  • Verbose and slower than a standard model, like other reasoning models
  • Reasoning traces consume context budget

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~20GBFits a 24GB card, with little room for long reasoning traces
Q5_K_M~23GBBetter quality, needs headroom beyond 24GB
Q8_0~35GBNear-lossless, needs 40GB or more
FP16~65GBFull precision, server or multi-GPU territory

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~20GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

For good quality

Q5_K_M · ~23GB needed

One 32GB GPU

NVIDIA GeForce RTX 5090

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M4 Pro .

64GB of system RAM alongside the card.

Best quality

FP16 · ~65GB needed

96GB of unified memory

AMD Ryzen AI Max+ 395 (Strix Halo)

or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .

Unified memory is shared with the model, so it is already counted above.

Licence

Apache 2.0 — read the licence

Benchmarks

BenchmarkScoreSourceAs of
AIME 202472.6% DeepSeek model card January 2025
GPQA Diamond62.1 DeepSeek-R1 technical report (Table 5) January 2025
MATH-50094.3 DeepSeek-R1 technical report (Table 5) January 2025
AIME 202556.3 Qwen3 technical report (re-evaluated baseline) May 2025

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

DeepSeek-R1-Distill-Qwen 32B: common questions

What hardware do I need to run DeepSeek-R1-Distill-Qwen 32B?
At its most compressed (Q4_K_M) it needs roughly 20GB of VRAM, and about 24GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is DeepSeek-R1-Distill-Qwen 32B free for commercial use?
Yes. DeepSeek-R1-Distill-Qwen 32B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run DeepSeek-R1-Distill-Qwen 32B on Apple Silicon?
Yes. DeepSeek-R1-Distill-Qwen 32B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does DeepSeek-R1-Distill-Qwen 32B run on CPU?
Yes, DeepSeek-R1-Distill-Qwen 32B can run on the CPU, though generation is slower than on a GPU.
What is DeepSeek-R1-Distill-Qwen 32B's context window?
DeepSeek-R1-Distill-Qwen 32B has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

  • MLX community MLX 4-bit and 8-bit

    MLX quants for Apple Silicon, usually 4-bit and 8-bit.

Recommended for

  • Strong local reasoning on a single 24GB card
  • Maths and problem solving without a server
  • Users who want R1-style reasoning under a permissive licence

Related models

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.