SSignal86
机器之心
1 sourcesByteDance Seed Team Uncovers New Scaling Variable for Text-to-Image: Caption Information, Not Length
ByteDance Seed team found that the training performance of text-to-image models depends not on caption length but on the amount of image-bound information within. They proposed Structured Prompt, which improves both diffusability and promptability, achieving significant gains on complex composition, reasoning, and world knowledge generation tasks. This research opens a new direction for scaling text-to-image models.