Prompt-rewriting model claims 50-point gain on 30-second video

single source· 1 articles · confidence: medium · first seen 2026-09-23 20:00 UTC

What this means for you

Nothing to deploy: this is an arXiv preprint with no weights, API or licence mentioned. If you work on video generation, the claim worth testing is that a second large model should plan shots before generation, and that the gain grows with clip length.

WanPE, a 397-billion-parameter model that turns a written prompt into a shot-by-shot plan — camera moves, lighting, sound — raises human preference for video from the Wan3.0 generator by 10.66 to 18.84 points on clips of 5–15 seconds and by 50.86 points at 30 seconds, the authors report. It was trained on 1.05 million real videos and tuned with a reinforcement-learning method the team calls SC-GRPO, which rewards keeping the user's stated requirements intact across shots. Their benchmark, WanPEval, holds roughly 11,000 blind pairwise judgements; no independent evaluation is reported.

Key facts

  • ·WanPE is a 397-billion-parameter prompt enhancement model trained on 1.05 million real-world videos. source
  • ·Reported human preference gain over raw user prompts is 10.66–18.84 points at 5–15 seconds and 50.86 points at 30 seconds. source
  • ·WanPEval covers durations from 5 to 30 seconds and is supported by roughly 11,000 blind pairwise assessments. source
  • ·Tuning used Semantic-Consistency GRPO; the paper reports reverse construction outperformed forward rewriting in ablations. source
  • ·The authors state WanPE leads all evaluated commercial offerings at 5–15 seconds and is competitive with Seedance 2.5 at 30 seconds. source

What the sources say

Sources

The original reporting. Follow these — they did the work.

← the wire