Brand Video Intelligence · methodology
How AI video evaluation is scored.
Brand Video Intelligence keeps three questions separate: is the video technically coherent, did it follow the brief, and does it obey the brand’s production grammar?
The separation matters because the questions require different evidence. A flicker measurement should not be presented as a model opinion; a semantic judgment should not be presented as a deterministic measurement. The scorecard labels both.
Three evaluations, one scorecard
| Layer | Method | What it reads | What it returns |
|---|---|---|---|
| Quality | Deterministic signals | Subject and background consistency, temporal flicker, imaging quality, motion behavior | A score and pass, warn, fail, or error status; evidence where available |
| Prompt adherence | Model judgment against the supplied brief | Objects, actions, scene, spatial relationships, text, artifacts, face integrity, audio | Per-dimension score, status, and diagnosis |
| Brand adherence | Comparison with a versioned brand profile | Color, grade, motion energy, editing pace, composition, typography, brand-mark rules | Distance-aware scores, rule violations, timestamps, and corrective instructions |
A brand profile is versioned evidence
Reference videos are read into a reusable profile covering distributions and qualitative rules: dominant colors, grade, motion energy, cut frequency, composition, typography, and brand-mark treatment. The profile records the extractor version and reference set so a later evaluation can identify which definition of the brand it used.
Different dimensions use different comparisons. Editing pace is not compared in the same way as typography; a hard brand-mark rule is not averaged away by a strong color score. The output keeps dimension scores and rule violations visible beside the overall verdict.
The verdict is meant to be acted on
The final verdict returns an overall score, a badge, hard failures, and recommended corrections. Corrections are structured so a person can review them or a generation workflow can use them as a prompt delta. A single overall number never replaces the underlying diagnosis.
Where frame-level evidence exists, the scorecard includes timestamps. That lets a reviewer inspect the relevant moment instead of watching the entire video to find the failure again.
What the score does not prove
- A high score does not predict campaign performance, audience response, or commercial success.
- Prompt-adherence and qualitative brand checks involve model judgment and can be wrong.
- A brand profile reflects its reference set; weak or contradictory references produce a weaker definition.
- Scores across different profiles are not automatically comparable because tolerances and rules may differ.
- Automated evaluation does not replace legal, regulatory, accessibility, or final editorial review.
- The beta may revise dimensions, thresholds, and labels as the evaluation contract develops.