Brand Video Intelligence · methodology

How AI video evaluation is scored.

Brand Video Intelligence keeps three questions separate: is the video technically coherent, did it follow the brief, and does it obey the brand’s production grammar?

The separation matters because the questions require different evidence. A flicker measurement should not be presented as a model opinion; a semantic judgment should not be presented as a deterministic measurement. The scorecard labels both.

Tier 01 / Layers

Three evaluations, one scorecard

LayerMethodWhat it readsWhat it returns
QualityDeterministic signalsSubject and background consistency, temporal flicker, imaging quality, motion behaviorA score and pass, warn, fail, or error status; evidence where available
Prompt adherenceModel judgment against the supplied briefObjects, actions, scene, spatial relationships, text, artifacts, face integrity, audioPer-dimension score, status, and diagnosis
Brand adherenceComparison with a versioned brand profileColor, grade, motion energy, editing pace, composition, typography, brand-mark rulesDistance-aware scores, rule violations, timestamps, and corrective instructions
Tier 02 / Profile

A brand profile is versioned evidence

Reference videos are read into a reusable profile covering distributions and qualitative rules: dominant colors, grade, motion energy, cut frequency, composition, typography, and brand-mark treatment. The profile records the extractor version and reference set so a later evaluation can identify which definition of the brand it used.

Different dimensions use different comparisons. Editing pace is not compared in the same way as typography; a hard brand-mark rule is not averaged away by a strong color score. The output keeps dimension scores and rule violations visible beside the overall verdict.

The verdict is meant to be acted on

The final verdict returns an overall score, a badge, hard failures, and recommended corrections. Corrections are structured so a person can review them or a generation workflow can use them as a prompt delta. A single overall number never replaces the underlying diagnosis.

Where frame-level evidence exists, the scorecard includes timestamps. That lets a reviewer inspect the relevant moment instead of watching the entire video to find the failure again.

Tier 03 / Limits

What the score does not prove

  • A high score does not predict campaign performance, audience response, or commercial success.
  • Prompt-adherence and qualitative brand checks involve model judgment and can be wrong.
  • A brand profile reflects its reference set; weak or contradictory references produce a weaker definition.
  • Scores across different profiles are not automatically comparable because tolerances and rules may differ.
  • Automated evaluation does not replace legal, regulatory, accessibility, or final editorial review.
  • The beta may revise dimensions, thresholds, and labels as the evaluation contract develops.