豆包大模型团队:Seedream 4.0技术报告(英文版)(19页).pdf
1、Seedream 4.0:Toward Next-generation MultimodalImage GenerationByteDance SeedAbstractWe introduce Seedream 4.0,an efficient and high-performance multimodal image generation systemthat unifies text-to-image(T2I)synthesis,image editing,and multi-image composition within asingle framework.We develop a h
2、ighly efficient diffusion transformer with a powerful VAE whichalso can reduce the number of image tokens considerably.This allows for efficient training ofour model,and enables it to fast generate native high-resolution images(e.g.,1K-4K).Seedream4.0 is pretrained on billions of textimage pairs spa
3、nning diverse taxonomies and knowledge-centric concepts.Comprehensive data collection across hundreds of vertical scenarios,coupledwith optimized strategies,ensures stable and large-scale training,with strong generalization.Byincorporating a carefully fine-tuned VLM model,we perform multi-modal post
4、-training for trainingboth T2I and image editing tasks jointly.For inference acceleration,we integrate adversarialdistillation,distribution matching,and quantization,as well as speculative decoding.It achievesan inference time of up to 1.4 seconds for generating a 2K image(without a LLM/VLM as PEmod
5、el).Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art resultson both T2I and multimodal image editing.In particular,it demonstrates exceptional multimodalcapabilities in complex tasks,including precise image editing and in-context reasoning,and alsoallows for multi-imag
6、e reference,and can generate multiple output images.This extends traditionalT2I systems into an more interactive and multidimensional creative tool,pushing the boundary ofgenerative AI for both creativity and professional applications.Seedream 4.0 is now accessible onVolcano Engine.Official Page:htt





点击查看更多