ERNIE vs Vidu
E
ERNIE
Baidu
Baidu's ERNIE image generation, a text-to-image line spanning the older ERNIE-ViLG and the newer open-weight ERNIE-Image diffusion transformer. Strongest on Chinese-language prompts, backed by Baidu's multimodal research.
Strengths
- ✓Strong Chinese-language prompt understanding
- ✓Open-weight ERNIE-Image runs at a compact 8B parameters
- ✓Competitive quality among open text-to-image models
Weaknesses
- ✕English-prompt fidelity trails Chinese
- ✕Docs and ecosystem are China-centric
- ✕Older ERNIE-ViLG variant is large and dated
✕
Vidu
Shengshu
Shengshu's Vidu video generation family, built on a U-ViT diffusion architecture. Known for strong character consistency, reference-to-video, and fast generation, with wide availability via API and app.
Strengths
- ✓Strong subject and character consistency
- ✓Reference-to-video and multi-image input
- ✓Fast generation turnaround
Weaknesses
- ✕Photorealism trails top-tier rivals
- ✕Prompt adherence weaker on complex scenes
- ✕Smaller Western community