Tencent Hy4 Preview
Architecture, Deployment Paths, and Hardware BoundariesA source-based profile of Tencent’s Hy4 Preview MoE model, covering its 770B/49B-active architecture, one-million-token context, official vLLM and SGLang recipes, FP8 path, and documented limitations.
This page organizes first-party documentation and editorial analysis. ModelRun Lab has not independently reproduced the published performance or hardware claims.
Model overview
Hy4 Preview is a Mixture-of-Experts language model released by the Tencent Hy Team. The official model card describes it as an early flagship preview aimed at coding, office work, game development, scientific research, and long-context agentic tasks.
The model has 770 billion backbone parameters, with 49 billion activated per token. It also includes a native multi-token-prediction layer with 10 billion total parameters and 0.7 billion activated parameters. ModelRun Lab has not independently run the checkpoint, measured its memory use, or reproduced Tencent’s evaluation results.
Published architecture
Tencent lists the following backbone specifications:
- 770B total parameters
- 49B activated parameters per token
- 78 layers
- 6,144 hidden size
- Gated DeepSeek Sparse Attention with IndexCache
- 64 attention heads
- 256 routed experts and one shared expert
- Eight routed experts plus the shared expert activated per token
- Four residual streams using identity Hyper-Connections
- One-million-token context length
- 120,832-token vocabulary
These are provider-published specifications. The model card excludes the native MTP layer from its backbone table.
Official model variants
Tencent publishes two closely related checkpoints:
- Hy4 Preview: the instruct model
- Hy4 Preview-FP8: the official FP8-quantized instruct model
The official deployment examples use the FP8 checkpoint. This makes the FP8 repository part of the same deployment topic rather than a separate model page. Runtime and hardware compatibility still depend on the selected software version, GPU architecture, precision support, context length, and serving configuration.
Official deployment paths
Tencent recommends vLLM or SGLang for production serving and links to first-party recipes for both runtimes. The documented examples expose an OpenAI-compatible API and serve the model as hy4-preview.
The vLLM example uses:
- The vllm/vllm-openai:hy4-preview image
- The official FP8 checkpoint
- Tensor parallel size 8
- Native MTP speculative decoding
- FLASHMLA sparse attention
- Hy4 reasoning and tool-call parsers
The SGLang example likewise uses the FP8 checkpoint, tensor parallel size 8, automatic reasoning and tool parsers, and NEXTN speculative decoding. These are provider examples, not universal minimum configurations.
Reasoning controls
The official quickstart states that reasoning defaults to high. Tencent recommends temperature 0.9 and top-p 1.0. For direct responses, the model card documents a no_think reasoning-effort option through chat-template arguments.
These settings are provider recommendations. Production applications should evaluate latency, output quality, token use, and tool behavior with representative workloads.
Hardware boundary
The official recipes use tensor parallelism across eight GPUs, but the reviewed model card does not specify a universal GPU model or minimum per-device VRAM figure. Actual requirements depend on the full or FP8 checkpoint, runtime kernels, GPU architecture, context length, KV cache, batch size, speculative decoding, and serving overhead.
For that reason, this page does not translate the eight-way example into a guaranteed hardware minimum and does not claim that ModelRun Lab has tested the deployment.
Known limitations
Tencent explicitly labels this as an early Hy4 version. The provider reports two known behavior issues: spending longer than necessary reasoning through complex tasks and a tendency to over-verify its own work. These limitations may affect latency, token usage, and workflow design.
Provider-reported evaluations
The model card includes benchmark and internal expert-evaluation results across software engineering, office analysis, games, and research tasks. Those results are provider-reported. They should be assessed with the published methodology, prompts, tool configurations, comparison conditions, and the model’s preview status.
License boundary
The official repository identifies the license as Apache 2.0. This page reports the repository metadata without providing legal advice. Users remain responsible for reviewing the current license, checkpoint terms, dependencies, input rights, output use, and applicable policies.
Evidence boundaries
- Architecture, context length, runtime recipes, parameter recommendations, and known limitations come from Tencent’s official model card.
- Benchmark and internal evaluation results are provider-reported.
- No universal GPU or VRAM minimum is asserted.
- ModelRun Lab has not independently tested either checkpoint.