DeepSeek-R1-Distill-Llama-8B

DeepSeek-R1-Distill-Llama-8B

About DeepSeek-R1-Distill-Llama-8B

DeepSeek-R1-Distill-Llama-8B is a distilled model based on Llama-3.1-8B. The model was fine-tuned using samples generated by DeepSeek-R1 and demonstrates strong reasoning capabilities. It achieved notable results across various benchmarks, including 89.1% accuracy on MATH-500, 50.4% pass rate on AIME 2024, and a rating of 1205 on CodeForces, showing impressive mathematical and programming abilities for an 8B-scale model

Metadata

Create on

License

MIT

Provider

DeepSeek

Specification

State

Deprecated

Architecture

Calibrated

No

Mixture of Experts

No

Total Parameters

8B

Activated Parameters

Reasoning

No

Precision

FP8

Context length

33K

Max Tokens

Ready to accelerate your AI development?

Ready to accelerate your AI development?

Ready to accelerate your AI development?