Hy3 replaces closed frontier APIs for teams that want to self-host a capable model under a permissive license, not chase the absolute top reasoning score. It's built for cost-sensitive agentic coding and long-context document work, not Western enterprises that need SOC2 or HIPAA compliance out of the box.
Hy3 is Tencent's open-weight Hunyuan model, released July 2026, built as a Mixture-of-Experts model with a large share of its parameters active per forward pass and hybrid fast-and-slow reasoning. It targets developers who want a self-hostable, permissively licensed model for agentic coding, not a closed API-only frontier model.
Provider: Tencent · Family: Hunyuan
Context window: 256,000 tokens · Max output: 32,768
Input modalities: text, image · Output: text
About Hy3
Hy3 is Tencent's third-generation Hunyuan large language model, officially released in July 2026 following a preview launch in April 2026. It is a Mixture-of-Experts model with 295 billion total parameters and 21 billion active parameters per forward pass, supporting a context window of up to 256,000 tokens. The model implements a hybrid fast-and-slow thinking architecture, allowing it to dynamically allocate compute between quick responses and extended reasoning chains. Hy3 was developed after Tencent rebuilt its pre-training and reinforcement learning infrastructure starting January 2026, guided by three principles: well-rounded capabilities across reasoning, long-context, instruction-following, and tool use; authentic evaluation beyond standard benchmarks; and model-inference co-design for cost efficiency. The inference efficiency improvement comes from deep co-optimization between the MoE architecture and inference frameworks, including compute performance gains and advanced quantization algorithms. Since its preview release, Hy3 has been integrated across Tencent's product ecosystem: Yuanbao gained agent functions for complex tasks generating PPT, Word, Excel, PDF, and HTML files; ima gained structured reasoning and long-form writing improvements; CodeBuddy and WorkBuddy saw a jump in PPT generation success rate; Marvis gained better file editing, diagnostics, and multi-agent collaboration. Real-world usage shows Hy3 reliably powers long agent workflows across document processing, data analysis, knowledge retrieval, and MCP toolchain orchestration. Hy3 is open-sourced under the commercially friendly Apache 2.0 license, available on Hugging Face, ModelScope, GitHub, and GitCode from day one. It supports mainstream inference frameworks vLLM and SGLang for direct deployment. Third-party platforms including OpenRouter and several agent IDEs are progressively integrating Hy3. See the pricing FAQ for exact per-token rates and personal plan pricing. Benchmark highlights from the preview period: strong overall usability and agent capability scores, exceptional STEM reasoning on complex tasks, and a meaningful inference efficiency gain at comparable intelligence density. The model shows strong in-context learning, instruction following, coding, and agent capabilities, particularly in repository-level coding tasks and search-integrated workflows via MCP.
Pricing
Tencent Cloud TokenHub pricing (Hy3 preview): input ~$0.18/M, cached input ~$0.06/M, output ~$0.59/M tokens. Personal agent platform plans from ~$4.10/month. Local inference is free under the Apache license, though it requires your own GPU (A100/H100 class). No batch discount published yet.
Key Features
- Hybrid Fast and Slow Reasoning: Dynamic compute allocation shifts between quick responses in fast mode and extended chain-of-thought in slow mode, one of the few open-weight MoE models built this way.
- Inference Efficiency via Model-Inference Co-Design: The MoE architecture is co-optimized with vLLM and SGLang for a real token-generation speedup at the same intelligence density, measured on Artificial Analysis.
- Large Context Window with High Recall: Supports up to 256,000 tokens of context with verified long-context needle-in-haystack performance, enabling whole-repository code analysis and multi-document Q&A.
- Apache-Licensed Open Weights with Full Ecosystem Support: Weights, code, and training recipes are published on Hugging Face, ModelScope, GitHub, and GitCode, with native vLLM and SGLang support plus permitted commercial use, fine-tuning, and distillation.
- Native MCP and Agentic Tool Use: RL-trained on MCP toolchains, reliably executing long multi-agent workflows across GitHub, filesystem, web search, and code execution.
Pros
- Strong intelligence per dollar among open-weight flagships, combining a real efficiency gain with a permissive open license for a compelling self-hosted cost curve.
- The hybrid reasoning architecture is genuinely novel, using dynamic MoE routing between fast and slow expert paths rather than a single thinking-token trick.
- Proven at real scale inside Tencent's own product ecosystem, with heavy token growth and many enterprise integrations since preview.
- A full open-source stack ships together: weights, code, recipes, and quantization configs, with no closed-training ambiguity.
Cons
- Peak reasoning trails several closed frontier models on the Intelligence Index; it isn't the outright smartest model available.
- Documentation and community support lean Chinese-first, which creates friction for Western developers relying on English resources.
- Tencent Cloud's international onboarding adds KYC friction compared with OpenRouter, Anthropic, or OpenAI APIs.
- Vision capabilities exist but are unbenchmarked against dedicated multimodal flagships, and there's no native audio or video I/O.
Benchmarks
- math: 72
- mmlu: 88
- mmlu pro: 78
- aime 2025: 75
- arc agi 2: 18
- humaneval: 88
- live bench: 58
- lmarena elo: 1320
- gpqa diamond: 62
- lmarena rank: 12
- aider polyglot: 65
- swe bench verified: 55
- humanitys last exam: 15
- artificial analysis intelligence index: 52
- artificial analysis price blended per m: 0.85
- artificial analysis speed tokens per sec: 95
Frequently Asked Questions
How much does Hy3 cost per 1M tokens?
Open weights are free to run yourself under the Apache license, though you supply the GPU. Tencent Cloud's TokenHub preview pricing runs about $0.18 per 1M input tokens, $0.06 per 1M cached input, and $0.59 per 1M output tokens, plus personal agent plans starting around $4.10 a month.
How does Hy3 compare on benchmarks vs closed frontier models?
Hy3 scores around 52 on the Artificial Analysis Intelligence Index, behind Claude Opus 4.8 and GPT-5.5 on raw reasoning, but it claims a real inference efficiency gain from its MoE and vLLM/SGLang co-design that keeps serving costs down at similar output quality.
Is Hy3 open source or proprietary?
Hy3 is open-weight under a permissive Apache license, one of the more developer-friendly options available. Weights are downloadable from Hugging Face, ModelScope, GitHub, and GitCode, and commercial use, fine-tuning, and distillation are all explicitly permitted.
Does Hy3 train on user data?
Tencent's hosted API retains inputs and outputs for 30 days for abuse monitoring unless flagged, with a zero-retention option for enterprise MaaS customers. Fully local, open-weight inference never sends your data anywhere.
Who is Hy3 best for and who should avoid it?
Hy3 fits teams that want a self-hostable, cost-efficient model for agentic coding and long-context document analysis under a permissive license. Avoid it if you need Western compliance certifications out of the box, English-first support, or the single highest reasoning score on the market.