
Tencent
Text Generation
Hunyuan-A13B-Instruct
Hunyuan-A13B-Instruct 僅啟用其 80 B 參數中的 13 B,卻能在主流基準上匹敵更大的 LLMs。它提供混合推理:每次呼叫可切換為低延遲“快速”模式或高精度“慢速”模式。內建 256 K-token 上下文,允許它在不減低功效的情況下解析書籍長度的文件。代理技能為 BFCL-v3、τ-Bench 和 C3-Bench 領導力而調校,使其成為優秀的自主助手基礎。分組查詢注意力和多格式量化提供記憶體輕量、GPU 高效的推理,適合現實世界的部署,並具備內建多語言支持和堅固的安全對齊,適用於企業級應用。...
總上下文:
131K
最大輸出:
131K
輸入:
$
0.14
/ M Tokens
輸入:
$
text
/ M Tokens
輸出:
$
0.57
/ M Tokens

Tencent
Text Generation
Hy4-preview
Hy4 preview is Tencent Hy's new-generation productivity flagship model, built on Gated DeepSeek Sparse Attention with iHC residual design, with approximately 770B total and 49B activated parameters per token. It natively supports a 1M context window with up to 64k output and a native MTP layer for speculative decoding to boost generation throughput. It demonstrates strong comprehension, planning, and sustained execution on complex tasks, and natively supports tool calls, structured output, and context caching. Deeply optimized for Agent, Coding, and productivity scenarios, it suits long-horizon planning and code workflows....
總上下文:
1049K
最大輸出:
262K
輸入:
$
0.834
/ M Tokens
輸入:
$
text
/ M Tokens
輸出:
$
2.501
/ M Tokens

Tencent
Text Generation
Hy3
Built for real-world business scenarios, Hy3 features a 295B/21B active MoE architecture, native 256K context support, and three reasoning modes. It enhances coding, long-form comprehension, multi-turn dialogue, and agentic task execution, balancing reliability, efficiency, and cost across both high-frequency interactions and complex workflows....
總上下文:
262K
最大輸出:
262K
輸入:
$
0.132
/ M Tokens
輸入:
$
text
/ M Tokens
輸出:
$
0.528
/ M Tokens

