Table of Contents

Tencent Hy4 Preview is an open-weight Mixture-of-Experts model designed for long-context reasoning, software development, research, and agent workflows. Its 770-billion-parameter backbone activates 49 billion parameters per token and supports a context window of up to one million tokens.
Those specifications make Hy4 Preview relevant for repository-scale coding, multi-file analysis, extended tool use, and large document collections. However, the “Preview” label matters: Tencent identifies over-reasoning and unnecessary self-verification as known limitations. Teams should evaluate the model on representative tasks before using it in production.
What Is Tencent Hy4 Preview?
Hy4 Preview is a text-generation model developed by the Tencent Hy Team. It uses a sparse MoE architecture to combine a very large total parameter count with a smaller active compute footprint for each generated token.
Tencent has released both the standard Hy4 Preview checkpoint and an FP8 version under the Apache License 2.0. The model can be self-hosted through supported inference frameworks such as vLLM and SGLang. Its official Hy4 Preview model card includes model weights, deployment instructions, recommended generation settings, and known limitations.
Hy4 Preview Core Specifications
Specification | Hy4 Preview |
|---|---|
Architecture | Mixture-of-Experts |
Backbone parameters | 770B |
Activated backbone parameters | 49B per token |
Backbone layers | 78 |
Context length | 1M tokens |
Routed experts per MoE layer | 256 |
Shared experts per MoE layer | 1 |
Routed experts activated per token | 8 |
Attention type | Gated DSA |
Residual streams | 4 |
Native MTP layer | 10B total, 0.7B activated |
License | Apache 2.0 |
Input modality | Text |
The figures for the 770B total and 49B activated parameters cover the backbone. The additional native Multi-Token Prediction layer contains 10B parameters, bringing the complete architecture to approximately 780B parameters.
Total vs. Activated Parameters
Total parameters describe the full capacity stored in the model. Activated parameters indicate how much of that capacity participates in processing each token.
Hy4 Preview contains 256 routed experts and one shared expert in each of its 77 MoE layers. For every token, the model selects eight routed experts while also using the shared expert. This routing allows the model to maintain broad specialization without computing all 770B backbone parameters for every token.
The 49B active figure should not be interpreted as the model’s deployment size. Hosting still requires storing the complete checkpoint, along with memory for the key-value cache, runtime operations, and long-context requests.
How Does Hy4 Preview Handle 1M-Token Context?
Hy4 Preview supports a stated context length of one million tokens. Its architecture uses sparse attention and cross-layer index reuse to reduce some of the computation associated with processing very long sequences.
A one-million-token limit is a capacity ceiling, not a guarantee of perfect recall across the entire prompt. Retrieval quality still depends on document structure, prompt design, information density, and where relevant evidence appears.
Gated DSA and IndexCache
Hy4 Preview uses Gated DeepSeek Sparse Attention, or Gated DSA. Instead of performing full attention across every token pair at every layer, the sparse indexer selects up to 2,048 relevant tokens for each query position.
IndexCache reduces additional overhead by allowing selected layers to reuse sparse token indices calculated by nearby layers. The underlying IndexCache research reports that, on a separate 30B DSA model, removing 75% of indexer computations produced up to a 1.82× prefill speedup and a 1.48× decoding speedup with limited quality loss.
Those performance figures come from the IndexCache research setup and should not be treated as guaranteed Hy4 Preview API latency.
iHC and Native MTP
The model also uses identity Hyper-Connections, or iHC, with four residual streams. According to Tencent’s architecture description, this design expands the paths through which information can move between layers.
Its native Multi-Token Prediction layer supports speculative decoding. The MTP layer proposes multiple future tokens, which the main model can verify together. This can reduce interactive latency when the inference engine, hardware, and workload are configured to benefit from speculative decoding.
MTP does not automatically improve every deployment. Under heavily batched workloads, the cost of generating and verifying draft tokens may offset the latency benefit.
What Can Hy4 Preview Do?
Hy4 Preview is best suited to tasks that combine long input, multi-step reasoning, code generation, or tool use. Smaller models may remain more economical for short classification, extraction, and routine chat requests.
Coding and Agent Workflows
The model is designed for long-horizon software tasks such as:
Exploring large repositories before proposing a change
Tracing behavior across source files, tests, and documentation
Planning and implementing multi-file refactors
Writing and repairing tests
Debugging failures through terminal or development tools
For reliable agent use, applications should still define tool permissions, execution timeouts, retry limits, and human approval points. A strong coding benchmark score does not make unrestricted shell access or automatic production deployment safe.
Document and Research Tasks
The large context window also supports document-heavy workflows, including contract comparison, technical literature review, financial analysis, and synthesis across multiple reports.
Developers should structure large inputs rather than placing an unorganized document collection into one prompt. Useful practices include adding document identifiers, asking the model to cite source sections, separating extraction from synthesis, and validating important conclusions against the original material.
Long context can reduce the number of retrieval steps, but it does not remove the need for retrieval systems. RAG remains useful when collections change frequently, exceed the context limit, require access control, or need traceable source selection.
How Does Hy4 Preview Perform?
Tencent reports substantial improvements over Hy3 across coding, agent, search, office, and reasoning evaluations. The results are promising, although most currently available figures are vendor-reported rather than independent measurements.
Coding and Agent Benchmarks
Selected results from Tencent’s benchmark appendix include:
Benchmark | Hy4 Preview Score | Task Focus |
|---|---|---|
Terminal-Bench 2.1 | 85.4 | Terminal-based agent tasks |
DeepSWE | 64.3 | Software engineering |
SWE Atlas Refactoring | 53.3 | Codebase refactoring |
OneMillionBench With Tools | 65.4 | Long-context tool use |
Toolathlon-Verified | 74.1 | Multi-tool agent tasks |
APEX-Agents, Pass@1 | 37.1 | Real-world professional agents |
Tencent evaluated models at their highest available reasoning settings and used benchmark-specific agent harnesses. Some comparison results were reproduced internally rather than taken from identical public evaluations.
The practical conclusion is that Hy4 Preview is worth testing for repository-scale coding and tool-based workflows. Production selection should still use a private evaluation set that reflects the team’s languages, repositories, tool definitions, latency targets, and acceptance criteria.
Known Preview Limitations
Tencent identifies two current behavioral issues: the model may reason longer than necessary on complex tasks and may over-verify its own output.
These tendencies can increase response time, output length, and inference cost. They may also cause an agent to repeat checks without materially improving the result.
Teams can manage this behavior by:
Setting explicit stopping conditions
Limiting tool calls and agent turns
Using shorter reasoning modes for direct tasks
Defining output schemas
Measuring cost per accepted result rather than cost per request
Requiring approval before irreversible actions
The model’s preview status also means prompts and evaluation thresholds may need to be revisited as later checkpoints are released.
How to Access and Use Hy4 Preview
Tencent provides official deployment recipes for vLLM and SGLang. Because the full checkpoint is exceptionally large, self-hosting requires substantial accelerator memory and distributed inference infrastructure. The FP8 checkpoint reduces the storage and memory burden, but it remains a large-scale deployment.
The official vLLM Hy4 Preview recipe configures sparse attention, MTP speculative decoding, tool-call parsing, and an OpenAI-compatible local endpoint.
Python API Example
After deploying Hy4 Preview through vLLM or SGLang, the local server can be called with the OpenAI Python client:
Tencent recommends temperature=0.9 and top_p=1.0 for the released checkpoint. The example uses the self-hosted model name and local endpoint documented in Tencent’s deployment instructions. A hosted service may use a different endpoint, model ID, context limit, or parameter set.
Hy4 Preview Pricing and Cost Calculation
Hy4 Preview pricing is not one universal rate. Costs depend on whether the model is self-hosted or accessed through a managed API.
For a hosted API, estimate the request cost with:
For example, a request containing 700,000 uncached input tokens, 200,000 cached tokens, and 50,000 output tokens would cost:
Use cached-input pricing only when the selected provider explicitly supports it for Hy4 Preview.
Self-hosting requires a different calculation. Relevant costs include GPU hours, cluster utilization, model loading time, long-context memory usage, networking, monitoring, and engineering support. The Apache 2.0 license removes a model-weight licensing fee, but it does not make inference free.
For either deployment model, compare cost per accepted task. A cheaper request can become more expensive if repeated tool calls, excessive reasoning, or failed outputs require additional runs.
Common Questions About Tencent Hy4 Preview 1M context, coding, deployment, and API costs
Q1. Is Hy4 Preview Open Source?
Yes. Tencent has released the Hy4 Preview and Hy4 Preview FP8 weights under the Apache License 2.0. Developers should still review the license and operational requirements before modifying, redistributing, or deploying the model commercially.
Q2. Does Hy4 Preview Support Image Input?
No. The released Hy4 Preview checkpoint is documented as a text-generation model, and its official deployment examples use text messages. Image input is not listed as a supported modality for this release.
Q3. Can 1M Context Replace Retrieval?
No. A 1M-token context window can reduce retrieval steps for bounded document sets, but RAG remains valuable for changing collections, source permissions, citations, lower prompt costs, and datasets that exceed the available context.
Q4. What Is Hy4 Preview FP8?
Hy4 Preview FP8 is Tencent’s lower-precision version of the instruct model. It is intended to reduce deployment memory and improve inference practicality on supported hardware while preserving the same general model architecture and one-million-token context capability.
Hy4 Preview is a strong candidate for long-context coding, research, and agent evaluation. Its open weights, sparse-attention design, and official serving recipes make experimentation possible, while its scale and preview-stage behavior require careful infrastructure planning, cost controls, and task-specific testing.
