Back to feed
News Story
SSignal86
机器之心
1 sources

ByteDance Seed Team Uncovers New Scaling Variable for Text-to-Image: Caption Information, Not Length

ByteDance Seed team found that the training performance of text-to-image models depends not on caption length but on the amount of image-bound information within. They proposed Structured Prompt, which improves both diffusability and promptability, achieving significant gains on complex composition, reasoning, and world knowledge generation tasks. This research opens a new direction for scaling text-to-image models.

Primary report

机器之心

Primary source