정보에 대해서DeepSeek-R1-Distill-Llama-8B
DeepSeek-R1-Distill-Llama-8B는 Llama-3.1-8B를 기반으로 하는 증류된 Model입니다. 이 Model은 DeepSeek-R1이 생성한 샘플을 사용해 미세 조정되었으며, 강력한 추론 능력을 보여줍니다. 이 Model은 다양한 벤치마크에서 주목할 만한 결과를 달성했으며, MATH-500에서 89.1%의 정확도, AIME 2024에서 50.4%의 합격률, CodeForces에서 1205점의 평가를 받으며 8B 규모 Model로서 인상적인 수학 및 프로그래밍 능력을 보여줍니다.
메타데이터
사양
주
Deprecated
건축
교정된
아니요
전문가의 혼합
아니요
총 매개변수
8B
활성화된 매개변수
추론
아니요
Precision
FP8
콘텍스트 길이
33K
Max Tokens
다른 모델과 비교
이 Model이 다른 것들과 어떻게 비교되는지 보세요.
DeepSeek
chat
DeepSeek-V3.2
Total Context:
164K
Max output:
164K
Input:
$
0.27
/ M Tokens
Output:
$
0.42
/ M Tokens
DeepSeek
chat
DeepSeek-V3.2-Exp
Total Context:
164K
Max output:
164K
Input:
$
0.27
/ M Tokens
Output:
$
0.41
/ M Tokens
DeepSeek
chat
DeepSeek-V3.1-Terminus
Total Context:
164K
Max output:
164K
Input:
$
0.27
/ M Tokens
Output:
$
1.0
/ M Tokens
DeepSeek
chat
DeepSeek-V3.1
Total Context:
164K
Max output:
164K
Input:
$
0.27
/ M Tokens
Output:
$
1.0
/ M Tokens
DeepSeek
chat
DeepSeek-V3
Total Context:
164K
Max output:
164K
Input:
$
0.25
/ M Tokens
Output:
$
1.0
/ M Tokens
DeepSeek
chat
DeepSeek-R1
Total Context:
164K
Max output:
164K
Input:
$
0.5
/ M Tokens
Output:
$
2.18
/ M Tokens
DeepSeek
chat
DeepSeek-R1-Distill-Qwen-32B
Total Context:
131K
Max output:
131K
Input:
$
0.18
/ M Tokens
Output:
$
0.18
/ M Tokens
DeepSeek
chat
DeepSeek-R1-Distill-Qwen-14B
Total Context:
131K
Max output:
131K
Input:
$
0.1
/ M Tokens
Output:
$
0.1
/ M Tokens
DeepSeek
chat
DeepSeek-R1-Distill-Qwen-7B
Total Context:
33K
Max output:
16K
Input:
$
0.05
/ M Tokens
Output:
$
0.05
/ M Tokens
