Vidu, developed by ShengShu Technology, a Tsinghua University-linked lab, and launched in 2024, generates video from text or image prompts, but built its reputation on Multiple-Entity Consistency, described as the first capability of its kind: blending unrelated reference images, people, objects, environments, into one video while keeping each element true to its original appearance.
That consistency mechanism directly addresses what the AI-video field calls character collapse, a character's appearance subtly drifting between separately generated shots, a persistent problem for anyone producing a series of related clips featuring the same character or product, with the current model supporting up to seven reference images per generation.
A free tier covers standard use, and paid plans start around $8 a month for more generations. For a creator producing a series of related clips featuring the same character, product, or setting, who needs that appearance to stay consistent across every shot, Vidu's Multiple-Entity Consistency feature addresses that specific, well-documented problem directly.






