Tencent Hy4 Preview: 1M Context, Coding, and API Guide

Table of Contents

Tencent Hy4 Preview is an open-weight Mixture-of-Experts model designed for long-context reasoning, software development, research, and agent workflows. Its 770-billion-parameter backbone activates 49 billion parameters per token and supports a context window of up to one million tokens.

Those specifications make Hy4 Preview relevant for repository-scale coding, multi-file analysis, extended tool use, and large document collections. However, the “Preview” label matters: Tencent identifies over-reasoning and unnecessary self-verification as known limitations. Teams should evaluate the model on representative tasks before using it in production.

What Is Tencent Hy4 Preview?

Hy4 Preview is a text-generation model developed by the Tencent Hy Team. It uses a sparse MoE architecture to combine a very large total parameter count with a smaller active compute footprint for each generated token.

Tencent has released both the standard Hy4 Preview checkpoint and an FP8 version under the Apache License 2.0. The model can be self-hosted through supported inference frameworks such as vLLM and SGLang. Its official Hy4 Preview model card includes model weights, deployment instructions, recommended generation settings, and known limitations.

Hy4 Preview Core Specifications

Specification

Hy4 Preview

Architecture

Mixture-of-Experts

Backbone parameters

770B

Activated backbone parameters

49B per token

Backbone layers

78

Context length

1M tokens

Routed experts per MoE layer

256

Shared experts per MoE layer

1

Routed experts activated per token

8

Attention type

Gated DSA

Residual streams

4

Native MTP layer

10B total, 0.7B activated

License

Apache 2.0

Input modality

Text

The figures for the 770B total and 49B activated parameters cover the backbone. The additional native Multi-Token Prediction layer contains 10B parameters, bringing the complete architecture to approximately 780B parameters.

Total vs. Activated Parameters

Total parameters describe the full capacity stored in the model. Activated parameters indicate how much of that capacity participates in processing each token.

Hy4 Preview contains 256 routed experts and one shared expert in each of its 77 MoE layers. For every token, the model selects eight routed experts while also using the shared expert. This routing allows the model to maintain broad specialization without computing all 770B backbone parameters for every token.


Sparse mixture-of-experts routing in Tencent Hy4 Preview, with token paths selecting a small active subset from a large expert array

The 49B active figure should not be interpreted as the model’s deployment size. Hosting still requires storing the complete checkpoint, along with memory for the key-value cache, runtime operations, and long-context requests.

How Does Hy4 Preview Handle 1M-Token Context?

Hy4 Preview supports a stated context length of one million tokens. Its architecture uses sparse attention and cross-layer index reuse to reduce some of the computation associated with processing very long sequences.

A one-million-token limit is a capacity ceiling, not a guarantee of perfect recall across the entire prompt. Retrieval quality still depends on document structure, prompt design, information density, and where relevant evidence appears.

Gated DSA and IndexCache

Hy4 Preview uses Gated DeepSeek Sparse Attention, or Gated DSA. Instead of performing full attention across every token pair at every layer, the sparse indexer selects up to 2,048 relevant tokens for each query position.

IndexCache reduces additional overhead by allowing selected layers to reuse sparse token indices calculated by nearby layers. The underlying IndexCache research reports that, on a separate 30B DSA model, removing 75% of indexer computations produced up to a 1.82× prefill speedup and a 1.48× decoding speedup with limited quality loss.

Those performance figures come from the IndexCache research setup and should not be treated as guaranteed Hy4 Preview API latency.

iHC and Native MTP

The model also uses identity Hyper-Connections, or iHC, with four residual streams. According to Tencent’s architecture description, this design expands the paths through which information can move between layers.

Its native Multi-Token Prediction layer supports speculative decoding. The MTP layer proposes multiple future tokens, which the main model can verify together. This can reduce interactive latency when the inference engine, hardware, and workload are configured to benefit from speculative decoding.

MTP does not automatically improve every deployment. Under heavily batched workloads, the cost of generating and verifying draft tokens may offset the latency benefit.


Hy4 Preview long-context sparse attention with selected tokens, reused indices across layers, and multi-token prediction

What Can Hy4 Preview Do?

Hy4 Preview is best suited to tasks that combine long input, multi-step reasoning, code generation, or tool use. Smaller models may remain more economical for short classification, extraction, and routine chat requests.

Coding and Agent Workflows

The model is designed for long-horizon software tasks such as:

  • Exploring large repositories before proposing a change

  • Tracing behavior across source files, tests, and documentation

  • Planning and implementing multi-file refactors

  • Writing and repairing tests

  • Debugging failures through terminal or development tools

  • Running iterative research and coding agents

For reliable agent use, applications should still define tool permissions, execution timeouts, retry limits, and human approval points. A strong coding benchmark score does not make unrestricted shell access or automatic production deployment safe.

Document and Research Tasks

The large context window also supports document-heavy workflows, including contract comparison, technical literature review, financial analysis, and synthesis across multiple reports.

Developers should structure large inputs rather than placing an unorganized document collection into one prompt. Useful practices include adding document identifiers, asking the model to cite source sections, separating extraction from synthesis, and validating important conclusions against the original material.

Long context can reduce the number of retrieval steps, but it does not remove the need for retrieval systems. RAG remains useful when collections change frequently, exceed the context limit, require access control, or need traceable source selection.


Controlled long-horizon coding and research agent workflow connecting repositories, tests, tools, documents, and human approval

How Does Hy4 Preview Perform?

Tencent reports substantial improvements over Hy3 across coding, agent, search, office, and reasoning evaluations. The results are promising, although most currently available figures are vendor-reported rather than independent measurements.

Coding and Agent Benchmarks

Selected results from Tencent’s benchmark appendix include:

Benchmark

Hy4 Preview Score

Task Focus

Terminal-Bench 2.1

85.4

Terminal-based agent tasks

DeepSWE

64.3

Software engineering

SWE Atlas Refactoring

53.3

Codebase refactoring

OneMillionBench With Tools

65.4

Long-context tool use

Toolathlon-Verified

74.1

Multi-tool agent tasks

APEX-Agents, Pass@1

37.1

Real-world professional agents

Tencent evaluated models at their highest available reasoning settings and used benchmark-specific agent harnesses. Some comparison results were reproduced internally rather than taken from identical public evaluations.

The practical conclusion is that Hy4 Preview is worth testing for repository-scale coding and tool-based workflows. Production selection should still use a private evaluation set that reflects the team’s languages, repositories, tool definitions, latency targets, and acceptance criteria.

Known Preview Limitations

Tencent identifies two current behavioral issues: the model may reason longer than necessary on complex tasks and may over-verify its own output.

These tendencies can increase response time, output length, and inference cost. They may also cause an agent to repeat checks without materially improving the result.

Teams can manage this behavior by:

  • Setting explicit stopping conditions

  • Limiting tool calls and agent turns

  • Using shorter reasoning modes for direct tasks

  • Defining output schemas

  • Measuring cost per accepted result rather than cost per request

  • Requiring approval before irreversible actions

The model’s preview status also means prompts and evaluation thresholds may need to be revisited as later checkpoints are released.


Evaluation controls for a preview AI model covering benchmarks, latency, cost, tool limits, stopping conditions, and human approval

How to Access and Use Hy4 Preview

Tencent provides official deployment recipes for vLLM and SGLang. Because the full checkpoint is exceptionally large, self-hosting requires substantial accelerator memory and distributed inference infrastructure. The FP8 checkpoint reduces the storage and memory burden, but it remains a large-scale deployment.

The official vLLM Hy4 Preview recipe configures sparse attention, MTP speculative decoding, tool-call parsing, and an OpenAI-compatible local endpoint.

Python API Example

After deploying Hy4 Preview through vLLM or SGLang, the local server can be called with the OpenAI Python client:

import os

from openai import APIError, OpenAI

client = OpenAI(
    base_url=os.getenv("HY4_BASE_URL", "http://127.0.0.1:8000/v1"),
    api_key=os.getenv("HY4_API_KEY", "EMPTY"),
    timeout=120.0,
)

try:
    response = client.chat.completions.create(
        model=os.getenv("HY4_MODEL", "hy4-preview"),
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a software engineering assistant. "
                    "Explain assumptions and identify files that require review."
                ),
            },
            {
                "role": "user",
                "content": (
                    "Review this repository change plan and identify missing "
                    "tests, migration risks, and rollback steps."
                ),
            },
        ],
        temperature=0.9,
        top_p=1.0,
    )

    print(response.choices[0].message.content)

except APIError as error:
    print(f"Hy4 Preview API request failed: {error}")
import os

from openai import APIError, OpenAI

client = OpenAI(
    base_url=os.getenv("HY4_BASE_URL", "http://127.0.0.1:8000/v1"),
    api_key=os.getenv("HY4_API_KEY", "EMPTY"),
    timeout=120.0,
)

try:
    response = client.chat.completions.create(
        model=os.getenv("HY4_MODEL", "hy4-preview"),
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a software engineering assistant. "
                    "Explain assumptions and identify files that require review."
                ),
            },
            {
                "role": "user",
                "content": (
                    "Review this repository change plan and identify missing "
                    "tests, migration risks, and rollback steps."
                ),
            },
        ],
        temperature=0.9,
        top_p=1.0,
    )

    print(response.choices[0].message.content)

except APIError as error:
    print(f"Hy4 Preview API request failed: {error}")
import os

from openai import APIError, OpenAI

client = OpenAI(
    base_url=os.getenv("HY4_BASE_URL", "http://127.0.0.1:8000/v1"),
    api_key=os.getenv("HY4_API_KEY", "EMPTY"),
    timeout=120.0,
)

try:
    response = client.chat.completions.create(
        model=os.getenv("HY4_MODEL", "hy4-preview"),
        messages=[
            {
                "role": "system",
                "content": (
                    "You are a software engineering assistant. "
                    "Explain assumptions and identify files that require review."
                ),
            },
            {
                "role": "user",
                "content": (
                    "Review this repository change plan and identify missing "
                    "tests, migration risks, and rollback steps."
                ),
            },
        ],
        temperature=0.9,
        top_p=1.0,
    )

    print(response.choices[0].message.content)

except APIError as error:
    print(f"Hy4 Preview API request failed: {error}")

Tencent recommends temperature=0.9 and top_p=1.0 for the released checkpoint. The example uses the self-hosted model name and local endpoint documented in Tencent’s deployment instructions. A hosted service may use a different endpoint, model ID, context limit, or parameter set.

Hy4 Preview Pricing and Cost Calculation

Hy4 Preview pricing is not one universal rate. Costs depend on whether the model is self-hosted or accessed through a managed API.

For a hosted API, estimate the request cost with:

Request cost =
(uncached input tokens ÷ 1,000,000 × input rate)
+ (cached input tokens ÷ 1,000,000 × cached-input rate)
+ (output tokens ÷ 1,000,000 × output rate)
Request cost =
(uncached input tokens ÷ 1,000,000 × input rate)
+ (cached input tokens ÷ 1,000,000 × cached-input rate)
+ (output tokens ÷ 1,000,000 × output rate)
Request cost =
(uncached input tokens ÷ 1,000,000 × input rate)
+ (cached input tokens ÷ 1,000,000 × cached-input rate)
+ (output tokens ÷ 1,000,000 × output rate)

For example, a request containing 700,000 uncached input tokens, 200,000 cached tokens, and 50,000 output tokens would cost:

0.70 × input rate
+ 0.20 × cached-input rate
+ 0.05 × output rate
0.70 × input rate
+ 0.20 × cached-input rate
+ 0.05 × output rate
0.70 × input rate
+ 0.20 × cached-input rate
+ 0.05 × output rate

Use cached-input pricing only when the selected provider explicitly supports it for Hy4 Preview.

Self-hosting requires a different calculation. Relevant costs include GPU hours, cluster utilization, model loading time, long-context memory usage, networking, monitoring, and engineering support. The Apache 2.0 license removes a model-weight licensing fee, but it does not make inference free.

For either deployment model, compare cost per accepted task. A cheaper request can become more expensive if repeated tool calls, excessive reasoning, or failed outputs require additional runs.

Common Questions About Tencent Hy4 Preview 1M context, coding, deployment, and API costs

Q1. Is Hy4 Preview Open Source?

Yes. Tencent has released the Hy4 Preview and Hy4 Preview FP8 weights under the Apache License 2.0. Developers should still review the license and operational requirements before modifying, redistributing, or deploying the model commercially.

Q2. Does Hy4 Preview Support Image Input?

No. The released Hy4 Preview checkpoint is documented as a text-generation model, and its official deployment examples use text messages. Image input is not listed as a supported modality for this release.

Q3. Can 1M Context Replace Retrieval?

No. A 1M-token context window can reduce retrieval steps for bounded document sets, but RAG remains valuable for changing collections, source permissions, citations, lower prompt costs, and datasets that exceed the available context.

Q4. What Is Hy4 Preview FP8?

Hy4 Preview FP8 is Tencent’s lower-precision version of the instruct model. It is intended to reduce deployment memory and improve inference practicality on supported hardware while preserving the same general model architecture and one-million-token context capability.

Hy4 Preview is a strong candidate for long-context coding, research, and agent evaluation. Its open weights, sparse-attention design, and official serving recipes make experimentation possible, while its scale and preview-stage behavior require careful infrastructure planning, cost controls, and task-specific testing.

Ready to accelerate your AI development?

Ready to accelerate your AI development?

Ready to accelerate your AI development?