Real run · self-run demonstration
Which model version is better — recraftv4 vs recraftv4_1.
A real evaluation we ran ourselves: five prompts, both model versions, scored on four dimensions. This is the deliverable format exactly as a client would receive it — on real generated images, not mock numbers.
ⓘ
Honest method. 10 images (5 prompts × 2 versions) generated at 1024×1024 with Recraft, then scored 0–100 per dimension by an automated visual judge. A production MetronLabs evaluation adds a blind human panel, inter-rater agreement and confidence intervals — the numbers below are single-judge and directional.
Result — score by dimension
| Dimension | recraftv4 (A) | recraftv4_1 (B) | Δ |
|---|---|---|---|
| Prompt adherence | 86.8 | 87.6 | +0.8 |
| Anatomy | 85.5 | 87.5 | +2.0 |
| Text rendering | 75.0 | 79.5 | +4.5 |
| Aesthetics | 83.6 | 87.8 | +4.2 |
Verdict: recraftv4_1 improves on every dimension. Biggest gains in aesthetics (+4.2) and text rendering (+4.5); prompt adherence is essentially flat (+0.8, within noise). If you're upgrading, v4.1 is the safe move — the one thing to keep watching is prompt adherence, where the gap is not yet meaningful.
The images that were scored

Prompts & what each tests
- P1 · Portrait — anatomy, skin, natural light.
- P2 · "MERIDIAN COFFEE" sign — text rendering (both spelled it correctly).
- P3 · 3 apples + 2 pears — prompt adherence / counting (both got the count right).
- P4 · Hands pouring latte — anatomy of hands.
- P5 · Rainy Tokyo alley — aesthetics and scene coherence.
Want this run on your own model or agent?
Emailhello@metronlabs.org
MoreAll documents