Hy4 preview 2026: Tencent 770B Apache weights and the blind eval
Tencent opened Hy4 preview. Separate the 770B/49B SKU, the 2.99/4 internal blind, and the preview caveats.
In early September 2026 Tencent’s Hunyuan Hy Team released Hy4 preview: a 770B-total, ~49B-active MoE with a 1M context window and Apache 2.0 weights. The story is not another parameter count. It is the lab’s own internal blind eval, and the preview defects it listed in public.
The query is Hy4 preview and 770B / 49B. This page follows Tencent’s announcement, the Hugging Face card tencent/Hy4-preview, and the ThursdAI September tracker. Some recaps date the first drop to 28 August; do not invent a single ship date the official posts do not print. Another open fleet that week: K2 Horizon.
What the lab said it shipped
The flagship is a sparse MoE: 770B total, about 49B active per token. The backbone has 78 layers — a dense FFN on layer 1, then 77 MoE layers with 256 routed experts plus one shared expert, top-8 routed plus the shared expert per token. A native MTP layer (10B total, ~0.7B active) is built in for speculative decoding. Attention is Gated DSA with IndexCache; residuals use iHC across four streams. Vocabulary 120,832; context 1M.
How to read the three numbers
- 1
Name the SKU before the average
Serving recipes name the vLLM image
vllm/vllm-openai:hy4-previewor SGLanglmsysorg/sglang:hy4-preview, with FP8 and 8-way tensor parallel in the sample. Reasoning defaults to high; passno_thinkfor a direct reply. Tencent’s announcement lists API prices at $0.834 / $2.501 / $0.042 per million tokens (input / output / cache hit). WorkBuddy and CodeBuddy were offered free for two weeks after launch; Hy3 free access was extended to 30 September. - 2
2.99 is an in-house blind rating
163 Tencent experts scored 203 engineering tasks. Hy4 preview averaged 2.99 / 4.00. Versus GLM-5.3: 2.92 (46.8% wins / 12.8% ties / 40.4% losses). Versus Kimi K3: 2.94 (51.2% / 7.9% / 40.9%). The card and the tencent.com post match. Judges were employees; the rubric is the lab’s. That is not a regulator finding and not a third-party rerun.
- 3
Quant numbers need a repo stamp
The official card points at AngelSlim. ThursdAI and the AngelSlim GGUF card describe an STQ1_0 build at about 2.38 bits per weight and about 214GB, versus roughly 1.5TB unquantized. That is a community quant recipe, not the Apache weights shrinking by themselves. Do not merge it with K2’s AA index. Another open-weight page: Qwen3.8-27B.
Discount the “recursive self-improvement” claim
The announcement says the model helped optimize training methods, data strategy, eval harnesses, and kernels, and that inference throughput rose 31.8% versus a baseline. There is no third-party reproduction invoice. Treat it as lab narrative, not an independent systems paper.
Hugging Face eval rows are not an AA index
The model page has hosted rows such as Apex Agents and Terminal-Bench. This article does not treat those rows as a verified third-party index. The checkable facts remain 770B/49B, Apache 2.0, and the 2.99/4 internal blind.
The preview lists its own defects
Tencent writes that reasoning chains run long and that the model over-verifies, with headroom left in pre-training and post-training. It frames the drop like Hy3 preview: ship early, hear what breaks. That is not a GA flagship claim.
| Claim | Checkable source | How to read it |
|---|---|---|
| 770B / 49B MoE, 1M context | HF card · GitHub Tencent-Hunyuan/Hy4-preview | Architecture from the official table |
| Blind 2.99 vs 2.92 / 2.94 | Tencent post · card “Built for Productivity” | Internal experts, not an independent rerun |
| Apache 2.0 weights | Model-card License | Read the repo LICENSE before download |
Tencent: this is an early Hy4. There is real headroom in pre-training and post-training. We would rather ship early and hear what breaks — that is what made Hy3 preview better.
# Hy4 preview · Tencent Hy Team · as of 2026-09
sku: 770B / 49B active # MoE, 1M context
license: Apache-2.0
layers: 78 # 1 dense + 77 MoE
experts: 256 routed + 1 shared; top-8
mtp: 10B / 0.7B active # speculative decode
blind_avg: 2.99/4 # 163 experts, 203 tasks
vs_glm53: 2.92 # 46.8 / 12.8 / 40.4
vs_kimi_k3: 2.94 # 51.2 / 7.9 / 40.9
api_usd_per_m: 0.834 / 2.501 / 0.042Boundaries
- Internal blind ≠ third-party index
- 2.99 / 2.92 / 2.94 are scores from Tencent experts. Do not relabel them as an AA Intelligence Index, and do not crown them against the same week’s closed-lab vendor tables.
- Apache 2.0 ≠ a 770B model on one box
- The official sample is 8-way tensor parallel plus FP8. AngelSlim’s low-bit GGUF is a different serving path; measure speed and quality on your own load.
- This page is not a meeting manual
- It only explains the release and the evals. For a short huddle with a board, create a room on wbmeet; see how to create and join a room.
Questions worth checking
Is Hy4 preview a new lab’s first weight drop?
No. It comes from Tencent’s Hunyuan / Hy Team. The announcement also names co-design with CodeBuddy, WorkBuddy, Yuanbao, and ima, plus a Hy3 free-access extension. Hy4 is this generation’s preview, not a founding date.
Does 2.99 mean it has “beaten” GLM-5.3 and Kimi K3?
Not by itself. The average is only 0.05–0.07 higher; the GLM-5.3 win rate is 46.8% with a 40.4% loss rate. Tencent’s own wording is “slightly ahead.” Judges were internal.
Is the 214GB quant the official default?
The official card ships AngelSlim tooling and an FP8 variant. The ~214GB / 2.38 bpw figure appears on the AngelSlim GGUF note and in weekly recaps. Use the filename you actually pull.
Does this prove open weights have matched closed flagships?
No. This page only unpacks the checkable release and the internal blind wording. Closed comparators need their own system cards and third-party indexes.