Meta · Llama 3 · based on meta-llama/Llama-3.1-70B
Llama 3.3 70B Instruct
Meta's 70B open-weight model with strong reasoning and 128K context; needs roughly 45 GB at INT4 plus KV cache.
916,029 downloads · 3,044 likes · 7,692 GitHub stars · Model card · GitHub · synced 2026-09-19
Specifications
- Parameters
- 70.6B
- Architecture
- dense
- Layers
- 80
- Hidden size
- 8,192
- KV heads · head dim
- 8 · 128
- Vocabulary
- —
- Native dtype
- —
- Context window
- 128,000 tokens
- Released
- 2024-11-26
- Last modified on Hub
- 2024-12-21
Features & licensing
- License
- Llama 3.3 Community
- Commercial use
- Yes
- Modalities
- text
- Tool / function calling
- No
- Reasoning mode
- No
- Languages
- EN, FR, DE, HI, IT, PT, ES, TH
- Quantised variants on Hub
- AWQ, EXL2, FP8, GGUF, bitsandbytes
- Pipeline
- text-generation
Task fit
Editorial scores (0–100) used by the recommendation engine.
- Chat90
- RAG / Q&A91
- Code84
- Summarisation90
- Extraction86
- Agents85
Hardware to run Llama 3.3 70B Instruct
Weights need about 155 GB at FP16, 82 GB at INT8 and 47 GB at INT4. The KV cache adds roughly 328 MB per 1,000 tokens per request. For a reference workload of 5 requests per second with 1,500 input and 300 output tokens, total GPU memory is around 81.4 GB, which fits on 2 × NVIDIA A100 (80 GB).
| Configuration | VRAM | Utilisation | Est. first token | Cloud / month |
|---|---|---|---|---|
| 4 × NVIDIA L4 (24 GB) | 96 GB | 85 % | ~0.5 s | $2,336 |
| 2 × NVIDIA L40S (48 GB) | 96 GB | 85 % | ~0.2 s | $2,774 |
| 2 × NVIDIA A100 (80 GB)recommended | 160 GB | 51 % | ~0.1 s | $4,672 |
| 2 × NVIDIA H100 (80 GB) | 160 GB | 51 % | ~0.0 s | $6,570 |
Estimates only. Run an assessment for your own traffic.
Llama 3.3 70B is a frequent 'highest quality' pick among open models. The community license permits commercial use with a monthly-active-user threshold that large enterprises should review.