
Tencent
Text Generation
Hunyuan-A13B-Instruct
Hunyuan-A13B-Instructは、その80 Bのパラメーターのうち13 Bのみをアクティブにしますが、主流のベンチマークでより大きなLLMに匹敵します。ハイブリッド推論を提供し、低遅延の「高速」モードまたは高Precisionの「低速」モードを各呼び出しごとに切り替えることができます。ネイティブの256 K-tokenコンテキストにより、劣化せずに本のような長さのドキュメントを処理できます。エージェントスキルはBFCL-v3、τ-Bench、C3-Benchのリーダーシップに合わせて調整されており、優れた自律型アシスタントのバックボーンとなっています。グループ化されたQuery Attentionと多形式の量子化により、メモリ効率の良い、GPUに優しいInferenceを実現し、実際の展開での使用に備えています。企業向けアプリケーションのためのマルチリンガルサポートと強固な安全性調整を備えています。...
Total Context:
131K
Max output:
131K
Input:
$
0.14
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
0.57
/ M Tokens

Tencent
Text Generation
Hy4-preview
Hy4 preview is Tencent Hy's new-generation productivity flagship model, built on Gated DeepSeek Sparse Attention with iHC residual design, with approximately 770B total and 49B activated parameters per token. It natively supports a 1M context window with up to 64k output and a native MTP layer for speculative decoding to boost generation throughput. It demonstrates strong comprehension, planning, and sustained execution on complex tasks, and natively supports tool calls, structured output, and context caching. Deeply optimized for Agent, Coding, and productivity scenarios, it suits long-horizon planning and code workflows....
Total Context:
1049K
Max output:
262K
Input:
$
0.834
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
2.501
/ M Tokens

Tencent
Text Generation
Hy3
Built for real-world business scenarios, Hy3 features a 295B/21B active MoE architecture, native 256K context support, and three reasoning modes. It enhances coding, long-form comprehension, multi-turn dialogue, and agentic task execution, balancing reliability, efficiency, and cost across both high-frequency interactions and complex workflows....
Total Context:
262K
Max output:
262K
Input:
$
0.132
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
0.528
/ M Tokens

