← All models

Bernini vs ERNIE

Bernini
Bernini
ByteDance

ByteDance's Bernini-R is a unified model for text-to-video and instruction-based video editing, from adding or removing objects to changing weather, background or art style. A semantic planner reads the instruction before a diffusion renderer produces the video.

Strengths

  • ✓One interface across generation and editing
  • ✓Strong instruction-following via a planning stage
  • ✓Open-source 1.3B renderer weights released

Weaknesses

  • ✕Open renderer is small, so fidelity trails larger models
  • ✕Very new with limited independent benchmarks
  • ✕Two-stage pipeline adds complexity
E
ERNIE
Baidu

Baidu's ERNIE image generation, a text-to-image line spanning the older ERNIE-ViLG and the newer open-weight ERNIE-Image diffusion transformer. Strongest on Chinese-language prompts, backed by Baidu's multimodal research.

Strengths

  • ✓Strong Chinese-language prompt understanding
  • ✓Open-weight ERNIE-Image runs at a compact 8B parameters
  • ✓Competitive quality among open text-to-image models

Weaknesses

  • ✕English-prompt fidelity trails Chinese
  • ✕Docs and ecosystem are China-centric
  • ✕Older ERNIE-ViLG variant is large and dated
See Bernini vs ERNIE in the full pricing comparison
Modelfal.ai
Bernini
Bernini R
★
CreditCrunch logo. a purple magnifying glass observing a gold coin

Compare for your exact needs

Set your budget, duration, and resolution.

81+ models · 333+ price points · from $3.99/week