DeepSeek-R1-Distill-Qwen-32B
Suverenita a hosting
Tento model je v katalógu len na referenciu. Zatiaľ nie je hostovaný na LLM EU a žiadny first-party endpoint ho tu neobsluhuje. Požiadavky naň idú priamo poskytovateľovi, podľa jeho vlastných podmienok.
Zhrnutie
DeepSeek-R1-Distill-Qwen-32B is a 32B dense model distilled from DeepSeek-R1 into a Qwen2.5 base, published under the MIT licence. It keeps the reasoning behaviour of R1 at a size that fits on one 80GB accelerator, with a configuration maximum of 131,072 tokens. DeepSeek documents it as a distillation, so its ceiling is the teacher model’s reasoning, not a new training run.
Špecifikácie
- Poskytovateľ
- DeepSeek (CN)
- Kontextové okno
- 131k
- Parametre
- 32.76B
- Modality
- text
- Licencia
- mit · licence
- Otvorené váhy
- yes
- Dátum vydania
- 20. 1. 2025
- Úlohy
- reasoning, coding, cheap, summarization
Jazyky
Silná: en, zh
Dostatočná: de, fr, es, it, pt, ru
Jazyky
Silná
- en
- zh
Dostatočná
- de
- fr
- es
- it
- pt
- ru
Neoverené: poskytovateľ nezdokumentoval toto tvrdenie. — a language is listed as strong only where a source claims it; the card links the evidence where one exists.
Kedy ho použiť
Silné stránky
- fits a single 80GB accelerator
- MIT licence
- retains the R1 reasoning style that smaller base models usually lack
Obmedzenia
- a distillation, so quality is bounded by the teacher
- long reasoning traces increase output cost
- no evaluation has been run by LLM EU