DeepSeek-R1-Distill-Qwen-32B

DeepSeek

Iba v katalógu Otvorené váhy

Suverenita a hosting

Iba v katalógu

Tento model je v katalógu len na referenciu. Zatiaľ nie je hostovaný na LLM EU a žiadny first-party endpoint ho tu neobsluhuje. Požiadavky naň idú priamo poskytovateľovi, podľa jeho vlastných podmienok.

Zhrnutie

DeepSeek-R1-Distill-Qwen-32B is a 32B dense model distilled from DeepSeek-R1 into a Qwen2.5 base, published under the MIT licence. It keeps the reasoning behaviour of R1 at a size that fits on one 80GB accelerator, with a configuration maximum of 131,072 tokens. DeepSeek documents it as a distillation, so its ceiling is the teacher model’s reasoning, not a new training run.

Špecifikácie

Poskytovateľ
DeepSeek (CN)
Kontextové okno
131k
Parametre
32.76B
Modality
text
Licencia
mit · licence
Otvorené váhy
yes
Dátum vydania
20. 1. 2025
Úlohy
reasoning, coding, cheap, summarization

Jazyky

Silná: en, zh

Dostatočná: de, fr, es, it, pt, ru

Jazyky

Silná

  • en
  • zh

Dostatočná

  • de
  • fr
  • es
  • it
  • pt
  • ru

Neoverené: poskytovateľ nezdokumentoval toto tvrdenie. — a language is listed as strong only where a source claims it; the card links the evidence where one exists.

Kedy ho použiť

Silné stránky

  • fits a single 80GB accelerator
  • MIT licence
  • retains the R1 reasoning style that smaller base models usually lack

Obmedzenia

  • a distillation, so quality is bounded by the teacher
  • long reasoning traces increase output cost
  • no evaluation has been run by LLM EU