Back to Model List

Hy Image3.5 preview – A High-Value Professional-Level Image Generation Model from Tencent HunYuan

AI Tech Editorial
RSS Feed
Hy Image3.5 preview – A High-Value Professional-Level Image Generation Model from Tencent HunYuan official screenshot
(Image source: official screenshot)

Executive Summary:

Hy Image3.5 preview is a high-value professional-level image generation model launched by Tencent HunYuan, designed to address the complex needs of high-quality image generation, precise text renderin...

1. What is Hy Image3.5 preview

Hy Image3.5 preview is a high-value professional-level image generation model launched by Tencent HunYuan, designed to address the complex needs of high-quality image generation, precise text rendering, and multi-round iterative editing in professional design and creative scenarios. The model supports both text-to-image and image-to-image modes, allows uploading up to 5 reference images per session for multi-subject fusion and stylized editing, and outputs images with a maximum resolution of 2K. It also supports continuous creation through conversational-style multi-round editing based on historical context. Compared to the previous generation, Hy Image3.0, its overall capabilities have improved by approximately 30%, with significant enhancements in text rendering, realism, and editing consistency. Internal blind testing at Tencent showed that its image generation quality is on par with Seedream 5.0 Pro, and it slightly outperforms Nano Banana Pro and Qwen-Image-3.0 Pro in certain dimensions. Through architectural simplification, decoupled deployment, and attention operator adaptation, the model reduces inference costs and offers API services to enterprise developers at a price lower than that of similar flagship products.

hy-image3-5-preview official website screenshot
Image source: Official article
Image source: official article

Technical positioning and domain: Hy Image3.5 preview belongs to the field of generative AI image models, positioning itself as a professional-level productivity tool. It covers application scenarios requiring high-precision visual output, such as commercial posters, film and television materials, e-commerce advertisements, UI design, and portrait editing. Its key differentiator lies in focusing on text rendering capabilities as the core entry point, directly addressing the common issue of text corruption in image generation with traditional diffusion models.

Development background: This model was developed by the Tencent HunYuan large model team, based on the design philosophy of the Hy Image3.0 series, with architectural simplification and systematic upgrades. The development motivation stems from Tencent's deep layout in AIGC productivity scenarios internally—ranging from the consumer-facing Ybao application to the ima knowledge management tool, and further to professional creative platforms such as Miora, WorkRally, and OnSolo. The HunYuan team has embedded image generation capabilities across multiple product lines in the Tencent ecosystem, achieving full coverage from consumer entertainment to professional creation.

Core value: This model solves three key issues in professional design scenarios: first, the accuracy and aesthetic quality of text rendering, enabling AI to directly generate images with extensive text elements such as posters, logos, and infographics without requiring post-processing; second, consistency in multi-subject editing, achieved through multi-reference image input and enhanced intent understanding, allowing complex operations such as outfit changes, local edits, and multi-subject fusion; third, cost control for high-resolution outputs, with technical optimizations reducing the API unit price for 2K images to 0.15 RMB per image, significantly lowering the threshold for enterprise-scale usage.

Technical features: The model employs a decoupled deployment architecture, separating the text branch and image branch for independent operation, optimization, and scheduling, thereby improving the efficiency of inference resource utilization. Combined with an attention operator adaptation plan optimized for underlying hardware, it effectively reduces memory usage and accelerates inference speed, providing the engineering foundation for low-cost, high-speed generation of 2K images. In terms of intent understanding, the model has undergone systematic upgrades centered on "from intent understanding to visual expression," enhancing its instruction-following and semantic comprehension capabilities, which significantly improves the precision and controllability of editing operations.

2. Key Features

  • Text-to-Image Generation: Generate high-quality images directly based on text prompts, covering various scenarios such as life documentation, commercial posters, UI design, and film materials. The model's mapping capability between semantic understanding and visual expression has significantly improved compared to the previous generation, accurately translating elements, styles, compositions, and other information from complex instructions into visual outputs.

  • Image-to-Image Generation: Supports uploading up to 5 reference images in a single session for operations such as changing outfits, local edits, style transfer, and multi-subject editing. The multi-reference image input mechanism enables the model to capture features of multiple subjects simultaneously, excelling in maintaining consistency across multiple subjects. This is ideal for creative tasks requiring the integration of multiple specified people or objects into a single image.

  • Multi-turn Dialogue Creation: Enables continuous generation and editing based on historical context, supporting progressive refinement. After users provide feedback on generated results, the model can understand changes in instructions within the context and maintain the continuity of existing elements. This is well-suited for design workflows that require repeated iterations and refinement, reducing the need for regenerating images from scratch and lowering associated costs.

  • Multi-format Output: Supports generating images in various aspect ratios and sizes, with a maximum output resolution of 2K. The flexible size adaptation capability allows the model to meet the publishing requirements of different platforms, from vertical formats for social media to horizontal formats for commercial posters, all achievable with a single click. The 2K resolution ensures detailed performance in large-scale printing or high-definition display scenarios.

  • Text Rendering: Significantly improved rendering and layout capabilities for small text, greatly reducing text corruption. This feature directly supports creative needs involving text elements such as posters, logos, and infographics. There have been notable improvements in details such as text box boundary control, font style matching, and multilingual text arrangement, driving an 80% increase in request volume for productivity scenarios.

  • Enhanced Editing Consistency: Instruction-following accuracy and multi-subject, stylized editing capabilities for image-to-image generation have been comprehensively improved, reducing bad case rates in advertising scenarios by 7%. Through upgraded intent understanding, the model can more accurately interpret users' localized modification requests, precisely altering specified areas while preserving the original appearance of unedited regions, thus minimizing unintended full-image changes.

3. How to Use

The integration methods for Hy Image3.5 preview cover three groups: end-users, knowledge management professionals, and enterprise developers. The specific usage paths are as follows:

  1. General Users (Yuanbao End): Open the Yuanbao app and navigate to the image generation feature section. Choose from preset templates such as portrait editing, intelligent photo editing, and free outfit swapping. Input a textual description or upload a reference image, set the desired aspect ratio, and click to generate. The generated results can directly enter a multi-round editing workflow, allowing users to adjust details based on historical context until satisfied, after which they can save the image. This path requires no technical configuration and is suitable for daily entertainment and light photo editing needs.

  2. Knowledge Management Users (ima End): Open the ima app and initiate image-related tasks via Copilot, including image pairing, posters, infographics, and PPT illustrations. Users describe their creative needs through dialogue or upload professional knowledge content as a reference. Hy Image3.5 preview converts textual or data information into visual outputs and supports iterative revisions based on feedback, delivering visual content tailored to knowledge management scenarios.

  3. Professional Creators (Miora/WorkRally/OnSolo): Access the corresponding professional creation platform (currently offering a two-week free trial). Select the image generation or editing task type. Upload up to 5 reference images and input creative instructions, such as scene descriptions or character consistency requirements. The model will generate asset images in bulk. Creators can iteratively refine the prompts based on the generated results to optimize the output. This is ideal for teams in comic studios, advertising agencies, and other environments requiring large-scale visual asset production.

  4. Enterprises and Developers (Tencent Cloud API): Access the model service via the Tencent Cloud API system. TokenHub can be used for token management, and integration with media processing (MPS), cloud video on demand (VOD), and other services can achieve a complete image generation and processing pipeline. Developers input text prompts and reference image URLs according to the API documentation parameters and call the generation interface to obtain results. Payment is based on the number of output images, with a unit price of 0.15 RMB per 2K image. Enterprises can integrate the generation capability into their own products to achieve process automation and large-scale invocation.

  5. Key Configuration Notes: When calling the API, choose appropriate aspect ratios and resolution parameters based on the scenario. The price and processing time for 2K output differ from standard resolution, so users should balance based on actual needs. When inputting multiple reference images, it is recommended to upload clear images with consistent lighting to achieve better fusion results. When performing multi-round edits, retaining the generation history as context can improve the accuracy of modification instructions.

  6. Best Practices: In text-intensive scenarios (e.g., posters), it is recommended to leverage the model's fine-text rendering capabilities by directly embedding the text into the prompt rather than adding it later. When using image-to-image generation for outfit swapping or multi-subject editing, the quantity and quality of reference images directly impact the consistency of the output. Enterprise users can connect the API with media processing services to create an integrated pipeline for generation, compression, and distribution.

4. Pros and Cons Analysis

Pros
Significant enhancement in comprehensive capabilities: Overall performance is improved by approximately 30% compared to Hy Image3.0. Internal blind testing at Tencent shows it is on par with Seedream 5.0 Pro and slightly better than Nano Banana Pro and Qwen-Image-3.0 Pro, making it competitive among flagship models.
Outstanding text rendering capability: There is a significant improvement in small text rendering and layout, greatly reducing text corruption issues. This has increased request volume in productivity scenarios such as posters, logos, and infographics by 80%, offering a differentiated advantage in Chinese text rendering.
Editing consistency and multi-subject capability: Instruction-following for image-to-image editing has been enhanced, supporting up to 5 reference images for multi-subject editing. This has reduced the bad case rate in advertising scenarios by 7%, making it suitable for complex editing needs such as e-commerce outfit changes and multi-character integration.
Cost-effectiveness and cost control: Decoupled deployment and attention operator adaptation reduce inference costs. The API unit price for 2K images is only 0.15 yuan per image, which is more cost-effective than the 4K output of similar flagship models ($0.24 per image), helping to lower the cost barrier for enterprise-scale usage.
Multi-turn conversational creation: Supports continuous generation and editing based on historical context, enabling progressive refinement and reducing the cost of repeated generation. It is suitable for design workflows requiring iterative refinement.

5. Comparative Analysis with Similar Tools

Comparison Dimension Hy Image3.5 preview (Tencent Hunyuan) Nano Banana Pro (Google) Seedream 5.0 Pro (ByteDance)
Positioning and Pricing High cost-performance professional image generation model, 2K image at 0.15 CNY per image Google flagship image generation model, 2K/4K options available, 4K at $0.24 per image ByteDance flagship image generation model, positioned for professional creation
Resolution Support Maximum 2K Supports 2K/4K output Maximum 2K (awaiting official confirmation for 4K support)
Reference Image Input Up to 5 reference images per request Multi-image fusion, maintains consistency of up to 5 main subjects Supports multi-reference image input
Text Rendering Significant improvement in small text rendering, suitable for posters, logos, infographics; productivity scenario request volume increased by 80% Industry-leading, supports long paragraphs and multi-language layout Strong Chinese text rendering capability
Editing Capabilities Local editing, style transfer, multi-subject editing; bad case rate in advertising scenarios reduced by 7% Fine-grained local editing, lighting/focus adjustment, camera transformation Strong consistency in image-to-image editing, supports local repainting
Blind Test Performance Slightly better than Nano-Banana Pro in internal GSB blind test (Tencent evaluation criteria) Slightly outperformed by Hy Image3.5 in Tencent's internal blind test On par with Hy Image3.5 (Tencent internal evaluation criteria)
Integration Channels Tencent Cloud API + Tencent products such as Yuanbao, ima, Miora Gemini API / Vertex AI / AI Studio Volcano Engine API + ByteDance products such as Jimo
Ecosystem Advantages Domestic compliance, RMB settlement, deep integration with Tencent office/creation tools Global ecosystem, supports real-time information grounding via search Synergy with ByteDance content ecosystem, optimized for short video scenarios

Selection Recommendations: In Chinese commercial design scenarios, Hy Image3.5 preview's text rendering capabilities and multi-reference image editing features make it a strong option for productivity needs such as posters, e-commerce materials, and film storyboards. It is particularly suitable for teams that have deeply integrated Tencent Cloud or Tencent office products. The 0.15 CNY per 2K image pricing provides a clear cost advantage in large-scale usage scenarios.

When ultra-high resolution output or global application development is required, Nano Banana Pro's 4K capabilities and integration with the Google Cloud ecosystem offer greater advantages, although teams must accept higher per-image costs and consider compliance issues for overseas API calls. For teams seeking full control over the model, the open-source weights of the Qwen-Image series provide a flexible path for private deployment, allowing for fine-tuning in local environments to adapt to specific styles or business needs.

6. Editor's Summary

The Hy Image3.5 preview demonstrates a clear path for Tencent Hunyuan's transition in the image generation domain from "general-purpose generation" to a "productivity tool." From a technological innovation perspective, the improvements in the model's architecture are noteworthy — a streamlined architecture combined with an efficient distillation algorithm maintains generation quality while compressing the model's size. The decoupled deployment strategy separates the text branch from the image branch, enabling independent optimization of the two subsystems, a design that holds significant engineering innovation in terms of inference resource utilization efficiency. The hardware adaptation plan for the attention operator also reflects the trend of model development extending from the algorithmic level to the system engineering level, laying the groundwork for large-scale, low-cost generation of 2K images.

From a practical value standpoint, the most direct contribution of Hy Image3.5 preview is elevating text rendering capabilities to a commercially viable level. The breakthrough in small text rendering and layout capabilities resolves the long-standing issue of garbled text in AI-generated images within scenarios involving text, and this single capability holds considerable real-world application value in fields such as advertising materials and brand design. Improvements in editorial consistency are also critical, with the quantitative data showing a 7% reduction in badcase rate indicating substantive progress in instruction-following and local editing accuracy, rather than just停留在演示层面 (停留在演示层面 is translated as "merely at the demonstration level"). The API pricing strategy of 0.15 yuan per image clearly reflects a cost-performance orientation, which is expected to promote the popularization of image generation in small and medium enterprises and high-frequency calling scenarios.

This model is primarily suitable for three user groups: design teams and advertising agencies that require mass production of visual materials containing text, C-end users seeking efficient portrait editing and entertainment-oriented features, and enterprise developers who wish to integrate image generation capabilities into their own products. For users already within the Tencent ecosystem, the deep integration of this model with products such as Yuanbao, ima, and Miora provides out-of-the-box convenience.

In future development, the competitiveness of Hy Image3.5 preview will depend on two directions: first, the addition of capabilities to output images at resolutions of 4K and above, to meet the demands of high-end printing and cinematic-level image quality, thus bridging the current gap in resolution specifications compared to Nano Banana Pro; second, enhancing the model's openness, such as through open-sourcing or offering private deployment options, which will help expand the user base beyond the Tencent ecosystem and build a broader developer community. In the current stage, the release of the official version of Hy Image3.5 preview and the publication of third-party benchmark evaluation data are worth continuous attention.

7. Application Scenarios

  • Commercial Posters and Brand Visuals: Thanks to the significant improvement in small text rendering capabilities, this model is well-suited for generating commercial posters, logo designs, and infographic materials with extensive textual content. The enhanced layout and text rendering greatly reduce the issue of text corruption commonly found in traditional AI image generation. Users can directly input brand titles, promotional messages, and event copy into the generation prompts to produce visual assets that are nearly ready for delivery, significantly reducing post-production correction costs. Productivity-related requests in this scenario have increased by 80%, validating the actual market demand.

  • Film and TV Materials and Storyboard Production: The WorkRally/OnSolo platform has already integrated the Hy Image3.5 preview for bulk visual asset generation in animation studios, including the creation of consistently styled asset images for props, characters, and scenes. The multi-reference image input mechanism supports cross-image consistency in character appearance, ensuring that the same character maintains a unified visual identity across multiple storyboards and reducing the "card-drawing rate" (i.e., the probability of generation results deviating from expectations), thereby improving the efficiency of bulk production. Creative teams can upload reference images centered around key characters to generate high-consistency, storyboard-level materials that match the description.

  • E-commerce Materials and Advertising Images: E-commerce operators can leverage the image-to-image generation feature to perform tasks such as clothing changes, product scene replacements, and batch generation of multi-angle display images. The badcase rate in advertising scenarios has decreased by 7% compared to the previous generation, indicating that the generated advertising materials are more reliable in terms of instruction-following and image stability, reducing the need for rework. Combined with the API batch calling mode, brands can quickly produce large volumes of visually consistent promotional materials during sales seasons, reducing the per-image production cost to as low as 0.15 RMB per image.

  • Portrait Editing and C-end Entertainment: The Yuanbao platform has already integrated features such as intelligent photo editing, free clothing changes, and photo-to-cartoon conversion. Ordinary users can directly upload their personal photos and use text instructions to complete style transformations or localized edits. The multi-round dialogue mechanism allows users to progressively modify the same image—first adjusting the background, then changing the filter, and finally switching the clothing style—without needing to re-upload the original image, ensuring a smooth and continuous experience. In this scenario, portrait editing-related requests have already seen a doubling in growth.

  • Office and Knowledge Visualization: Enterprise knowledge management users can use ima Copilot to convert document content, data reports, or professional knowledge into visual outputs such as illustrations, PPT slides, and knowledge infographics. The model understands structured information and translates it into visual expressions, enabling complex concepts to be presented in intuitive graphical formats. This is applicable to internal training materials, industry analysis reports, and academic presentations, lowering the barrier for non-designers to produce visual content.

8. FAQ

Q: What are the main improvements of Hy Image3.5 preview compared to Hy Image3.0 and earlier versions?
A: According to official data, the overall capabilities of Hy Image3.5 preview have improved by approximately 30% compared to the previous generation, with the enhancements concentrated in three dimensions: text rendering, photorealism, and editing consistency. Among these, improvements in small text rendering and layout capabilities are particularly noticeable in production scenarios, driving an 80% increase in related request volumes. The enhanced instruction-following for image-to-image generation has reduced the badcase rate in advertising scenarios by 7%, and the stability of multi-subject editing has also improved.

Q: The current version is labeled as "preview"—when will the official version be released?
A: According to official information, Hy Image3.5 preview is the pre-release version for this stage, and the timeline for the official version has not yet been disclosed. One of the main purposes of the preview version is to collect real feedback from developers and creators to iteratively optimize model performance. Functionality and pricing strategies across various channels may change in the official version. Users are advised to continue monitoring official announcements from Tencent Hunyuan for updates on the official release.

Q: How does 2K output compare to competitors' 4K capabilities in practical use?
A: Hy Image3.5 preview supports a maximum resolution of 2K (approximately 2048×2048 or equivalent aspect ratios), while Nano Banana Pro offers 4K output options. In practical applications, 2K resolution is sufficient for the majority of online display, social media posting, standard print, and high-resolution electronic screen display needs. The advantage of 4K mainly lies in scenarios involving ultra-large outdoor advertisements or large-format artistic printing. If the business scenario does not require ultra-large output, the 0.15 RMB per image pricing of Hy Image3.5 is more attractive for budget control. For customers who require native 4K output, Nano Banana Pro must be used currently, and the cost is relatively higher.

Q: What are the best practices for using multiple reference images? Are there any limitations on image uploads?
A: Up to 5 reference images can be uploaded in a single request for multi-subject fusion, outfit changes, or style editing. To achieve optimal results, it is recommended that the reference images have clear subjects, even lighting, and simple backgrounds, and should avoid occluded, blurry, or highly watermarked images. In multi-image inputs, the imaging size and angle of each subject should be as consistent as possible to improve the model's ability to capture consistent human features. The official documentation provides detailed information on image format, size, and file size limitations (e.g., common JPG/PNG formats, aspect ratio requirements, and a maximum file size of approximately 10MB). Developers should review the corresponding specifications before calling the API.

Q: What is the billing method and usage limits for the API service?
A: The API service is billed based on the number of output images, with a price of 0.15 RMB per 2K image. Specific pricing for standard resolutions (such as 1K or lower), the concurrent call limit per account, the request per minute (RPM) limit, and the number of reference images that can be uploaded are all detailed in the latest pricing and traffic limit pages for the corresponding services (such as TokenHub, Media Processing MPS, and Cloud Video On Demand VOD) on Tencent Cloud console or API documentation. The actual cost may vary depending on the deployment region, call peak, and the associated cloud services used. It is recommended to refer to the bill and platform announcements for accuracy.

Q: Does this model support precise local modifications to generated images, rather than regenerating the entire image?
A: Yes. The model provides local editing capabilities through its image-to-image mode. Users can annotate or describe the areas of the reference image that need modification (e.g., replacing the background, adjusting facial expressions, or modifying specific objects). The model maintains the content of non-edited areas and only modifies the target region. Editing consistency has improved significantly compared to the previous generation, with a 7% reduction in badcase rates in real production scenarios such as advertising materials, effectively reducing the common issue of "changing one part and messing up the whole image." The multi-round dialogue mechanism also supports multiple consecutive local adjustments to the same result, making it suitable for design processes with high detail requirements.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.