
Tencent
Text Generation
Hunyuan-A13B-Instruct
Hunyuan-A13B-Instruct仅激活其80 B参数中的13 B,但在主流基准测试中与更大的LLMs匹配。它提供混合推理:低延迟的“快速”模式或高精度的“慢速”模式,可以在每次调用时切换。本地256 K-token上下文让它能够处理书籍长度的文档而不退化。代理技能为BFCL-v3、τ-Bench和C3-Bench领导进行了调优,使其成为出色的自主助手骨干。分组查询注意力加上多格式量化提供记忆轻、GPU高效的推理,用于现实世界的部署,具有内置的多语言支持和企业级应用的强大安全对齐。...
上下文长度:
131K
最大输出长度:
131K
Input:
$
0.14
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
0.57
/ M Tokens

Tencent
Text Generation
Hy4-preview
Hy4 preview is Tencent Hy's new-generation productivity flagship model, built on Gated DeepSeek Sparse Attention with iHC residual design, with approximately 770B total and 49B activated parameters per token. It natively supports a 1M context window with up to 64k output and a native MTP layer for speculative decoding to boost generation throughput. It demonstrates strong comprehension, planning, and sustained execution on complex tasks, and natively supports tool calls, structured output, and context caching. Deeply optimized for Agent, Coding, and productivity scenarios, it suits long-horizon planning and code workflows....
上下文长度:
1049K
最大输出长度:
262K
Input:
$
0.834
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
2.501
/ M Tokens

Tencent
Text Generation
Hy3
Built for real-world business scenarios, Hy3 features a 295B/21B active MoE architecture, native 256K context support, and three reasoning modes. It enhances coding, long-form comprehension, multi-turn dialogue, and agentic task execution, balancing reliability, efficiency, and cost across both high-frequency interactions and complex workflows....
上下文长度:
262K
最大输出长度:
262K
Input:
$
0.132
/ M Tokens
Input:
$
text
/ M Tokens
Output:
$
0.528
/ M Tokens

