Want an ad-free experience? Upgrade your plan.
Z-Image is an advanced next-generation image generation foundation model. It is designed to deliver exceptionally fast, high-quality, and highly controllable image synthesis using a novel Single-Stream Diffusion Transformer (S3-DiT) architecture. Z-Image integrates all modalities—text, semantic tokens, and VAE image tokens—into one unified sequence, greatly improving parameter efficiency and generation speed.
With 6 billion parameters, Z-Image aims to set a new standard in the open-source ecosystem, offering performance comparable to or surpassing larger proprietary models while maintaining much faster inference.
Technical innovations that make Z-Image a leader in efficient generative AI
Z-Image-Turbo requires only 8 NFEs to generate an image—a massive improvement over traditional diffusion models. Sub-second generation on enterprise GPUs, real-time performance on consumer hardware.
Exceptional Chinese and English text rendering. Generate posters, ads, product packaging, and social media graphics with accurate typography in both languages.
Prompt Enhancer enables semantic and contextual interpretation. Better understanding of object relationships, accurate adherence to complex multi-step instructions, and coherent compositions.
Natural language-driven image-to-image transformations. Add or remove objects, change style or lighting, modify backgrounds—all with flexible, creative control.
The sweet spot of AI modeling. Large enough for deep comprehension, small enough for 16GB consumer GPUs. Single-Stream DiT processes text and image tokens together.
Fully open for commercial use, research, and community modification. No royalty fees. Fine-tune on your own datasets for custom styles.
Flexible ecosystem for high-speed generation and advanced creative editing
Only 8 NFEs for photorealistic results. Sub-second on enterprise GPUs, real-time on 16GB consumer hardware. Ideal for interactive apps and rapid prototyping.
Complete non-distilled model for researchers and power users. Full flexibility for custom domain adaptation and fine-tuning on specific datasets.
Natural language-driven image editing. Add/remove objects, change style, adjust lighting—maintains structural integrity while applying creative transformations.
Simple workflow for developers and artists
Choose Z-Image-Turbo for speed, Z-Image-Base for customization, or Z-Image-Edit for post-processing. All variants are open-source and ready for deployment.
Download pre-trained weights and run locally. Z-Image runs smoothly on 16GB VRAM GPUs. ComfyUI provides a modular node-based interface for limitless customization.
Z-Image natively understands and renders both English and Chinese. Input intricate descriptions or poetic phrases—Z-Image interprets your semantic intent with high precision.
With Z-Image-Turbo's sub-second latency, iterate through dozens of concepts rapidly. Use Z-Image-Edit to perfect details with natural language commands.
Recommended next steps based on what you're working on — pick up where this tool leaves off.

Upload a still and watch AI bring it to life with natural motion.
Turn still images into cinematic video
Open
Set two key frames — AI fills the smooth transition between them.
Turn still images into cinematic video
Open
Fine-tune images with AI editing
Open
Fix blur, noise & detail in one click — old photos look brand-new.
Enhance image quality before publishing
Open