Tencent Hunyuan dropped Hy4 preview on August 28, 2026, and did the launch in a way that nobody expected. No SWE-bench table. No Terminal-Bench chart. No long list of benchmark scores. The official X post was one sentence: "Go use it, and tell us what's broken." The model itself — 770B total parameters, 49B active, 1M context — was open-sourced under Apache 2.0 at the same moment it went live inside WorkBuddy, CodeBuddy, 元宝, and ima.
This is a different posture from Zhipu, which ten days earlier had published a full-page benchmark comparison against Fable 5 and GPT-5.6 Sol. Tencent's message is: benchmarks don't settle this anymore. Real users, real engineering tasks, real workflows do. So they shipped it as "preview," put it in front of their daily-active user base for a free two-week window, and asked for feedback.
The numbers that do exist are worth taking seriously — and worth reading with the skepticism the "preview" label invites.
What's New: 770B/49B MoE, 1M Context, Apache 2.0

Hy4 preview is the successor to Hy3, and the generational jump is large:
| Spec | Hy3 | Hy4 preview |
|---|---|---|
| Total parameters | 295B | 770B |
| Active parameters | 21B | 49B |
| Context window | 256K | 1M (1.024M tokens) |
| License | open | Apache 2.0 |
| Input price | ¥1/M | ¥6/M |
| Output price | ¥4/M | ¥18/M |
| Cache hit | — | ¥0.3/M |
That is 2.6× total parameters, 2.3× active parameters, and 4× context length in one generation. It ships as 170 weight files on HuggingFace and GitHub. The positioning word is "productivity" — not agents, not reasoning, but productivity. The four target workloads are software engineering, office/knowledge work, game development (Tencent's home turf), and research (AI R&D, molecular dynamics, condensed matter physics, basic math).
The model was co-designed with Tencent's own products. Tencent says the training data was built jointly with experts across software engineering, gaming, finance, and security, and that Hy4 was designed in tandem with CodeBuddy and WorkBuddy from the start — not a model built in a lab and then adapted to products, but a model trained on the real workflows those products run.
Pricing in international terms is roughly $0.834/M input and $2.501/M output. That is about 40% cheaper than GLM-5.3 (~$1.40 / ~$4.40), but about 6× more expensive than GLM-5.3-Flash. This is flagship-tier pricing, not Flash pricing. The ¥0.3/M cache-hit price is the lever for long-running coding agents that repeatedly reuse system prompts and retrieval context.
Architecture: MoE at 770B, Pretraining + Post-Training Together

Hy4 preview is a Mixture-of-Experts model. The 770B total / 49B active split means most of the model's parameters sleep for any given token; only ~49B of experts are routed in. That's how you get 1M context and 770B total without per-token compute exploding.
Tencent says both pretraining and post-training improved this generation — this is the deliberate contrast with GLM-5.3, which explicitly shipped on a frozen GLM-5.2 base with all gains from post-training. Tencent's claim is that they expanded model size, context length, and data scale on the pretraining side, and simultaneously pushed post-training on the productivity-data side. The result, they say, is "another major jump in intelligence."
The most unusual sentence in the press release, though, is this: Hy4 preview "participated in its own development process." Tencent says the model was involved in training methods, data strategy, evaluation frameworks, and automatic optimization of low-level operators. The model proposes approaches, runs experiments, and feeds code, logs, and feedback back into the next round of exploration. They call it "an early-stage recursive self-improvement loop."
The concrete claim attached to that is throughput: Hy4 preview autonomously analyzed its own inference-system bottlenecks and ran multiple rounds of operator fusion and communication optimization, raising end-to-end throughput by 31.8% over its baseline.
Two readings are required here. The generous reading: an AI-assisted infrastructure optimization loop that produced a real, measured throughput gain. The less generous reading: "recursive self-improvement" is one of the most loaded phrases in AI safety discourse, and Tencent used it in an official launch announcement. The 31.8% number is self-reported with no disclosed baseline and no third-party reproduction. Treat it as a direction-of-travel signal, not a spec sheet. Watch for whether a technical report later substantiates it.
Benchmarks: No Public Table, One Blind Test, One Third-Party Score

This is where Tencent's launch style is most distinctive. The official announcement contains no public benchmark table. The only performance comparison they lead with is an internal blind test:
- 163 internal expert reviewers
- 203 real engineering tasks
- Hy4 preview: 2.99 / 4.00
- Kimi K3: 2.94 / 4.00
- GLM-5.3: 2.92 / 4.00
Read that carefully. Hy4 wins, but by 0.05–0.07 points. Tencent's own wording is "slightly ahead." On a 4-point scale, that's a tie with a narrow edge, not a blowout. And the reviewers are all Tencent employees, the task set is private, and the calling configuration (thinking tier, tool budget, environment) is not disclosed. It is directional evidence at best.
The only hard third-party number comes from BenchLM's SWE-bench Pro leaderboard (verified August 28, covering 65 models):
| Rank | Model | Vendor | Weights | Score |
|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | closed | 80.3% |
| 2 | Claude Fable 5 | Anthropic | closed | 80.0% |
| 3 | Claude Opus 5 | Anthropic | closed | 79.2% |
| 6 | Qwen3.8 Max | Alibaba | open | 67.7% |
| 7 | Hy4 preview | Tencent | open | 65.7% |
| 8 | Ornith-1.5-397B | Ornith | open | 65.1% |
| 9 | Grok 4.5 | xAI | closed | 64.7% |
| 10 | GPT-5.6 Sol | OpenAI | closed | 64.6% |
Three takeaways: Hy4 is #2 among open-weight models, behind Qwen3.8 Max by 2 points and ahead of Grok 4.5 and GPT-5.6 Sol. The gap to the closed frontier (the Claude trio at ~80%) is still 14 points — open has not caught closed on coding yet. And the leaderboard itself carries the standard caveat: setups differ across vendors, and OpenAI's own July audit estimated ~30% of SWE-bench Pro tasks are flawed.
The absence of a benchmark table is itself a signal. By late 2026, public benchmark trust has eroded enough that a major vendor chose to lead with "use it for two weeks for free and tell us what's broken" instead of a chart. Zhipu's anonymous ox-alpha stunt and Tencent's literal "preview" naming are two sides of the same strategy: when benchmarks are suspect, real-user friction becomes the credibility mechanism.
What It Means for Developers

The two-week free window is the actual launch event. WorkBuddy and CodeBuddy are free for two weeks. Hy3's free period was extended to September 30. If you're evaluating Hy4, this is the window to run your hardest real engineering tasks against it — not read about it. Tencent explicitly framed the "preview" label as a feedback-collection exercise.
API access paths. Tencent Cloud TokenHub and OpenRouter both carry Hy4 preview. Pricing is ¥6/M input, ¥18/M output, ¥0.3/M cache hit. For long coding-agent runs that reuse a system prompt and tool definitions, the cache-hit tier is where the bill actually comes down.
Self-hosting expectations. 770B total parameters is not a consumer-machine model. Expect serious multi-GPU deployment or quantized runs; check the official repo for quantization guidance. The realistic self-host route is a mid-size datacenter cluster, not a workstation.
Tencent ecosystem integration. If you're already inside Tencent's product stack — CodeBuddy, WorkBuddy, 元宝, ima — Hy4 preview is the default upgrade path, and it's co-designed for those workflows out of the box. If you're not in that ecosystem, the API path still works, but you lose the product co-design advantage.
Think-mode granularity. Hy4 preview only exposes two thinking tiers: high and no_think. You can fully disable reasoning, but there's no intermediate "medium" tier. That's a constraint worth knowing before you build routing logic around it.
Comparison to GLM-5.3 and Kimi K3. On price, Hy4 is ~40% cheaper than GLM-5.3. On the internal blind test, it edges both by a hair. On SWE-bench Pro, it's #2 open at 65.7%. None of these say "switch everything today." They say "add Hy4 to your evaluation matrix and run it on your real workload." That's literally what Tencent is asking you to do.
The strategic read. Tencent has been open that Hunyuan is a capital-expenditure line item. Shipping a 770B Apache 2.0 model is effectively giving part of that capex away to the ecosystem in exchange for positioning. Combined with the "model participated in its own training" claim, the longer-term signal is that Tencent wants to be the open-platform default for productivity agents — the model you build on because it's free, Apache-licensed, and integrated with the largest product surface in Chinese consumer tech.
Bottom Line
Hunyuan Hy4 preview is a 770B/49B, 1M-context, Apache 2.0 flagship that ships without a benchmark table, leads with a 163-expert blind test it wins by 0.05 points, and sits at #2 on SWE-bench Pro among open weights at 65.7%. It is ~40% cheaper than GLM-5.3 and free for two weeks inside Tencent's own products. The "model helped build itself" claim and the 31.8% self-optimized throughput number are the parts to watch — promising but unverified. For developers, the right move is the one Tencent literally asked for: take the free two-week window, run your hardest real engineering task, and decide from evidence rather than from a press-release chart.
Resources
- Tencent Hunyuan Official — model page and release announcement
- Tencent Cloud TokenHub — API access and pricing
- HuggingFace: tencent/Hy4-preview — open weights under Apache 2.0
- Securities Star Report — launch coverage and internal blind test details
- Xinhua Report — availability, pricing, and product launch details
Hunyuan Hy4 preview is available via API at Tencent Cloud TokenHub and on OpenRouter. Weights are open-sourced on HuggingFace under Apache 2.0. Pricing: ¥6/M input, ¥18/M output, ¥0.3/M cache hit. Free two-week trial on WorkBuddy and CodeBuddy.