DeepSeek-R1-Distill-Qwen-32B

DeepSeek

Catalogue only Open weights

Sovereignty and hosting

Catalogue only

This model is catalogued for reference. It is not hosted on LLM EU yet, and no first-party endpoint serves it here. Requests to it go to the provider directly, under that provider's own terms.

Summary

DeepSeek-R1-Distill-Qwen-32B is a 32B dense model distilled from DeepSeek-R1 into a Qwen2.5 base, published under the MIT licence. It keeps the reasoning behaviour of R1 at a size that fits on one 80GB accelerator, with a configuration maximum of 131,072 tokens. DeepSeek documents it as a distillation, so its ceiling is the teacher model’s reasoning, not a new training run.

Specifications

Provider
DeepSeek (CN)
Context window
131k
Parameters
32.76B
Modalities
text
License
mit · licence
Open weights
yes
Release date
20 Jan 2025
Tasks
reasoning, coding, cheap, summarization

Languages

Strong: en, zh

Adequate: de, fr, es, it, pt, ru

Languages

Strong

  • en
  • zh

Adequate

  • de
  • fr
  • es
  • it
  • pt
  • ru

Not verified: the provider has not documented this claim. — a language is listed as strong only where a source claims it; the card links the evidence where one exists.

When to use it

Strengths

  • fits a single 80GB accelerator
  • MIT licence
  • retains the R1 reasoning style that smaller base models usually lack

Limits

  • a distillation, so quality is bounded by the teacher
  • long reasoning traces increase output cost
  • no evaluation has been run by LLM EU