Hy4-preview
About Hy4-preview
Hy4 preview is Tencent Hy's new-generation productivity flagship model, built on Gated DeepSeek Sparse Attention with iHC residual design, with approximately 770B total and 49B activated parameters per token. It natively supports a 1M context window with up to 64k output and a native MTP layer for speculative decoding to boost generation throughput. It demonstrates strong comprehension, planning, and sustained execution on complex tasks, and natively supports tool calls, structured output, and context caching. Deeply optimized for Agent, Coding, and productivity scenarios, it suits long-horizon planning and code workflows.
Available Serverless
Run queries immediately, pay only for usage
Input Price
$
0.834
/ M Tokens
Cache Read
$
0.042
/ M Tokens
Output Price
$
2.501
/ M Tokens
Metadata
Specification
State
Available
Architecture
Calibrated
No
Mixture of Experts
No
Total Parameters
770B
Activated Parameters
Reasoning
No
Precision
FP8
Context length
1049K
Max Tokens
262K
Supported Functionality
Serverless
Supported
Serverless LoRA
Not supported
Fine-tuning
Not supported
Embeddings
Not supported
Rerankers
Not supported
Support image input
Not supported
JSON Mode
Supported
Structured Outputs
Not supported
Tools
Supported
Fim Completion
Not supported
Chat Prefix Completion
Supported
Compare with Other Models
See how this model stacks up against others.

Tencent
chat
Hunyuan-MT-7B
Total Context:
33K
Max output:
33K
Input:
$
0.0
/ M Tokens
Output:
$
0.0
/ M Tokens

Tencent
chat
Hunyuan-A13B-Instruct
Total Context:
131K
Max output:
131K
Input:
$
0.14
/ M Tokens
Output:
$
0.57
/ M Tokens

Tencent
chat
Hy4-preview
Total Context:
1049K
Max output:
262K
Input:
$
0.834
/ M Tokens
Output:
$
2.501
/ M Tokens

Tencent
chat
Hy3
Total Context:
262K
Max output:
262K
Input:
$
0.132
/ M Tokens
Output:
$
0.528
/ M Tokens

Tencent
chat
Hy3-preview
Total Context:
262K
Max output:
Input:
$
0.066
/ M Tokens
Output:
$
0.26
/ M Tokens
