Llama 3.3 70B Instruct
Meta · Text generation · 70B · 131k context · Released 6 December 2024
Permitted with conditions
Open weights
Runs on CPU
Apple Silicon
Meta's 70-billion-parameter instruction-tuned model, delivering performance close to their much larger 405B model at a fraction of the hardware cost. A strong general-purpose choice if you have the VRAM for it.
Strengths
- Strong general reasoning and instruction following
- Long 128k context window
- Excellent ecosystem support across every major inference engine
Weaknesses
- Needs serious hardware, not a laptop model at usable quality
- Weaker at code than dedicated coding models of similar size
- Licence carries conditions that matter at scale
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~43GB | Good balance, the common choice for 48GB cards |
| Q5_K_M | ~50GB | Noticeably better than Q4, needs more headroom |
| Q8_0 | ~75GB | Near-lossless, but you need multiple GPUs |
| FP16 | ~141GB | Full precision, server territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Llama 3.3 Community License — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| MMLU | 86.0 | Meta model card | December 2024 |
Availability
- Official page
- Hugging Face
- ollama run llama3.3:70b
Recommended for
- General-purpose assistant on a 48GB+ setup
- Long-document analysis
- Mac Studio users with 64GB+ unified memory
Related guides
Glossary
Catalogue entry last verified 15 January 2026. Specifications change; verify anything you are about to spend money on.