Skip to content
TECHNICAL REPORT V1.0AUGUST 2026//SPEAR SYSTEMS RESEARCH

DORU-1:preview

A multi-modal creative evaluation engine for pre-execution ad analysis. Scores video and static creatives before a dollar of media spend is committed.

LIVEMEMORABILITY · SALIENCY · EMOTION · COMPLIANCE

> 01. Abstract

DORU-1:preview is a multi-modal creative evaluation engine that analyzes static and video advertisements before media capital is deployed. A single asset produces a structured dossier of diagnostic panels: visual saliency and attention curves, memorability scores, transcripts, emotion fingerprints, hook grades, copy and CTA analysis, weak-moment flags, advisory compliance screening, and deterministic audit receipts.

The core contribution is the memorability head: a compact neural predictor trained on human long-term recall data (LAMBDA, WACV 2025; 1,749 participants, 2,205 advertisements). On the LAMBDA test split it achieves Spearman ρ = 0.413 (static) and ρ = 0.345 (video) - 54% and 45% of the human split-half consistency ceiling (ρ = 0.77).

We additionally report the first published benchmark of 2026 frontier models on ad memorability: Gemini 3.6 Flash (0.268), GPT-5.6 Luna (0.247), Grok 4.5 (0.217), Sonnet 5 (0.144), and Gemini 3.1 Pro (0.131) all score below both DORU-1 heads on the same 219-ad protocol.

> 02. The Benchmark

DORU-1 does not predict engagement, click-through, or conversion. It predicts and measures the components that precede them: attention, memorability, emotional response, message clarity, and pacing. Activation is not attention. Attention is not engagement. Engagement is not purchase intent. Each layer is a separate problem, and conflating them is how creative-scoring products fail.

Protocol: identical to the published baselines for comparability. The LAMBDA 219-test split; the dataset's per-ad verbalizations (brand, title, per-scene description, emotions, tags, on-screen text, tone, colors); 10 stratified few-shot examples; a 0–100 scoring prompt; Spearman ρ vs ground truth. No vision, text-only, matching the literature protocol.

See how these measurements become ranked creative decisions in the public sample dossier.

LAMBDA test split - Spearman rank correlation

Modelρn95% CI
Human split-half consistency (ceiling)0.77--
DORU-1 static0.413214[0.29, 0.52]
DORU-1 video0.345214[0.22, 0.46]
Gemini 3.6 Flash (thinking)0.268219[0.140, 0.386]
GPT-5.6 Luna (high)0.247219[0.118, 0.367]
Grok 4.5 (fast)0.217219[0.087, 0.340]
Sonnet 5 (medium)0.144219[0.011, 0.271]
Gemini 3.1 Pro0.131219[−0.001, 0.259]
GPT-4o 10-shot (published)0.18219-
GPT-3.5 10-shot (published)0.06219-
Benchmark scope: DORU-1 is compared against general-purpose frontier models and the human consistency ceiling. Single-task, memorability-only research systems are out of scope given DORU-1's multi-panel evaluation surface.
Reading the confidence intervals: ρ is an estimate of the true correlation, and at ~214 samples that estimate carries uncertainty. A 95% confidence interval is the range in which the true value would fall 95 times out of 100 if the experiment were repeated. A wide interval means the measurement is imprecise, not that the result is weak. All intervals are wide because the test set is small, a property of the benchmark shared equally by every model. DORU-1 intervals are computed on the 214 test ads with complete feature extraction; frontier models are evaluated on the full 219.
FIGURE 01 - LAMBDA BENCHMARK: SPEARMAN RANK CORRELATION BY MODEL
● point estimate · [CI] 95% confidence interval · scale 0.00 to 0.80
0.000.200.400.600.80HUMAN CEILING0.77DORU-1 STATIC0.413DORU-1 VIDEO0.345GEMINI 3.6 FLASH0.268GPT-5.6 LUNA0.247GROK 4.50.217SONNET 50.144GEMINI 3.1 PRO0.131GPT-4o 10-SHOT0.18GPT-3.5 10-SHOT0.06

Ratio to the human consistency ceiling

FIGURE 02 - RATIO TO HUMAN CONSISTENCY CEILING (ρ / 0.77)
Full bar = human consistency ceiling. Gold = DORU-1 heads. Dim = 2026 frontier models, same protocol.
DORU-1 STATIC
0.54x
DORU-1 VIDEO
0.45x
GEMINI 3.6 FLASH (THINKING)
0.35x
GPT-5.6 LUNA (HIGH)
0.32x
GROK 4.5 (fast)
0.28x
SONNET 5 (med.)
0.19x
GEMINI 3.1 PRO
0.17x

The frontier plateau

Three generations of generalist models: 0.06 (GPT-3.5, 2023) to 0.18 (GPT-4o, 2025) to a 0.13–0.27 band across the 2026 frontier. The gap to the human ceiling has not moved. A dedicated head, trained on the right data, did.

FIGURE 03 - THE FRONTIER PLATEAU: THREE GENERATIONS, ZERO MOVEMENT
0.00.80.60.40.2HUMAN CEILING 0.77202320252026 FRONTIERDORU-1GPT-3.5 0.06GPT-4o 0.18GEMINI 3.1 PRO 0.131SONNET 5 0.144GPT-5.6 LUNA 0.247GROK 4.5 0.217GEMINI 3.6 FLASH 0.268FEB 19JUN 30JUL 9JUL 14JUL 21AUG 11DORU-1 STATIC 0.413DORU-1 VIDEO 0.3452026 frontier models ordered by release date, all below both DORU-1 heads in point estimate.Gap vs Gemini 3.6 Flash and GPT-5.6 Luna is within sampling noise at n≈214; the three-generation plateau is not.

Reading the table

  • All five 2026 frontier models score below both DORU-1 heads in point estimate, with DORU-1 using a compact dedicated regressor — orders of magnitude fewer parameters — against frontier models with ~100B+ active parameters.
  • Pairwise gaps vs DORU-1: the differences vs Grok 4.5, Sonnet 5, and Gemini 3.1 Pro are statistically robust; the gaps vs Gemini 3.6 Flash and GPT-5.6 Luna are directionally consistent but individually within sampling noise at n≈214. Larger test sets, a planned scale-up path, will tighten these intervals.
  • Why: memorability is an unintuitive task. As the TOT2MEM study (Bhattacharyya et al., WACV 2026) found, general multimodal reasoning and world knowledge may not be sufficient for predicting recall signals. A dedicated head trained on the right labels extracts signal generalists cannot.
  • Category position: DORU-1 is the only full-stack creative evaluation platform benchmarked on this protocol; no comparable multi-panel system exists to compare against. The comparison above is against general-purpose frontier models and the human consistency ceiling.
  • Protocol notes: Sonnet 5 compressed its output scale (15–65, 22 unique values), so its ρ is likely an underestimate. Gemini 3.1 Pro's result is borderline (p=0.052). All runs are single-shot; the literature baselines averaged 3 seeds. Grok 4.5's initial run fabricated sequential IDs and was discarded; the reported value is a clean re-run.

> 03. System Overview

DORU-1 is served within the Spear platform. Creatives are uploaded and analyzed asynchronously, and results are returned as structured dossiers that drive the buyer/publisher review exchange. All evaluation compute runs on Spear Systems infrastructure.

Analysis pipeline

FIGURE 06 - ANALYSIS PIPELINE
CREATIVESTATIC / VIDEOINGESTIONASYNCPARALLEL PANELSSALIENCYMEMORABILITYEMOTIONTRANSCRIPTHOOK GRADECOMPLIANCE+ AUDIT RECEIPTDOSSIERSTRUCTURED

Analysis panels (v1)

PanelOutput
Video saliency60-frame attention curve + heatmap
Static saliency256-cell attention heatmap
Memorability (video)score 0–1 + metrics
Memorability (static)score 0–1 + metrics, text-augmented
Transcriptword/segment timing
Emotion fingerprint4D vector, confidence-tagged
ABCD featureshook, pacing, CTA, faces, speech, supers
Hook gradePASS/FAIL
Copy/CTA analysisheadline/body/CTA assessment
Weak-moment flagstimestamped attention-dip flags
Compliance screeningadvisory flags, timecoded
Audit receiptdeterministic, replay-verifiable

The language analysis layer uses a fallback chain of hosted and self-hosted analysis models. Operational guardrails bound queue depth and per-partner throughput.

> 04. Memorability Heads

Architecture

The memorability heads are compact neural regressors layered on frozen, pre-trained visual backbones, trained on human long-term recall data. The video head fuses raw low-level features with scene-level and event-level representations; the static head is augmented with a language signal derived from the creative's copy, transcript, and on-screen text.

Training

Labels come from LAMBDA (Si et al., WACV 2025): long-term brand-recall scores from 1,749 participants across 2,205 real advertisements (276 brands, 113 industries, avg 33s). Split: 1,964 train / 219 test.

Trained with a standard regression objective on the train split, with a held-out validation split for checkpoint selection and raw-feature normalization fit on train only. Backbones remain frozen; only the compact regressor trains.

Results

HeadTest Spearman (n=214)% of ceiling (0.77)
Static (text-augmented)0.41354%
Video0.34545%
Ceiling: the human split-half consistency on LAMBDA brand recall is ρ = 0.77 (25 random split-half trials, paper §2.2). A perfect model cannot exceed it, because the labels themselves are noisy estimates.

> 05. Saliency and Attention

DORU-1 maps viewer attention second-by-second. Video assets produce a 60-frame attention curve and heatmap; static assets produce a 256-cell attention heatmap.

The attention curve is converted to a saliency-entropy trace; sustained entropy dips mark moments where visual interest collapses. These are surfaced as timestamped weak-moment flags: structural signals for editors, telling them where to insert a pattern interrupt or re-hook. The model identifies where; the creative team decides what.

FIGURE 04 - RESPONSE CURVE WITH WEAK-MOMENT DETECTION (SALIENCY TRACE)
0%100%0s5s10s15s20s25sWEAK MOMENT 8.8s-9.5s

The first-3-second window is evaluated against a rubric (face presence, bold supers, clear problem statement, auditory alignment) yielding a PASS/FAIL hook grade. This is a structural check, not a prediction of retention.

> 06. Language Analysis Layer

The language analysis layer produces the qualitative panels:

  • Emotion fingerprint: a 4D vector (trust, excitement, discomfort, memorability) with per-dimension confidence tags. Subcortical reward signals are explicitly confidence-tagged: they are proxies, not measurements.
  • ABCD creative features: hook, pacing, CTA, faces, speech, supers, inferred from frames and audio.
  • Copy/CTA analysis: headline, body, and call-to-action assessment against the creative's own claims.

The layer is decoupled from the memorability heads: swapping the analysis model does not change the memorability score, and vice versa.

FIGURE 05 - EMOTION FINGERPRINT (4D, CONFIDENCE-TAGGED)
TRUST 0.71 [hi]EXCITEMENT 0.64 [med]MEMORABILITY 0.47 [med]DISCOMFORT 0.22 [med]

> 07. Compliance Screening (Advisory)

DORU-1 scans transcripts and on-screen text for claim patterns (financial, medical, outcome-based) and surfaces timecoded flags into the review exchange.

The screening is advisory by design. A flagged asset enters the review queue, the publisher acknowledges with a comment, and the buyer retains final compliance authority. It accelerates the compliance workflow; it does not replace standard compliance checks.

> 08. Audit Receipts

Every analysis produces a deterministic, replay-verifiable receipt binding the asset, the model version, and the results. Any change to the analysis configuration invalidates prior receipts by construction. Scores are reproducible or they are not published.

DORU-1 ANALYSIS RECEIPT · REPLAY-VERIFIED
CREATIVE ID      : 8f3a-2c91-4e77-9b10
ENGINE           : doru-1:preview
MODEL VERSION    : v0.1.0
MEDIA HASH       : a1b2c3...f0e1
PANELS           : saliency · memorability · emotion
                   · transcript · hook · compliance
RECEIPT          : 9f86d081...0f00a08
VERIFICATION     : REPLAY VERIFIED - inputs and outputs
                   bound; any config change invalidates
STATUS           : APPROVED WITH CONDITIONS

> 09. Syntagma-1 - Roadmap

Syntagma-1:preview is the planned enterprise engine: a neural-response layer and a macro-context layer, gated on infrastructure readiness. Architecture details are available under NDA.

Neither layer is live in DORU-1. DORU-1 is explicitly creative-only; macro context is excluded by design.

CapabilityDORU-1:previewSyntagma-1
StatusLIVEIN DEVELOPMENT
Input modalitiesStatic + videoFull multi-modal
Memorability scoring✓✓
Saliency / attention✓✓
Compliance screeningadvisoryadvisory
Neural response layer-planned
Macro context layerexcluded by designplanned
Primary use caseHigh volumeHigh capital

> 10. Risk Factors

  • 01Scores are creative-only. DORU-1 evaluates the creative itself; seasonality, market, and broader context are not yet part of scoring.
  • 02Compliance screening is advisory. The buyer retains final compliance authority; screening accelerates review but does not gate.
  • 03Norm benchmarks are still building. Percentile baselines strengthen as more creatives run through the system.

> 11. Licensing and Corpus Strategy

All model components are commercially licensed for use, verified at the weights level, not just the repository license. The benchmark dataset (LAMBDA) is publicly released under a permissive license (labels MIT-licensed; source footage remains under the original creators' licenses). No non-commercial or research-only weights are used in the shipped engine.

Partner data usage is governed by platform terms agreed at onboarding.

> 12. Conclusion

DORU-1:preview is a pre-execution creative evaluation engine with real, measured, reproducible results: memorability heads at 54% and 45% of the human consistency ceiling, and the first published benchmark showing 2026 frontier models scoring below both heads on the same protocol.

It is not the finished product. Norm baselines await corpus volume, and the macro layer is roadmap. The architecture, the measurements, and the limitations are documented so the claims can be checked, not just believed.

> 13. References

  • Si et al., "Long-Term Ad Memorability: Understanding & Generating Memorable Ads," WACV 2025 (arXiv 2309.00378) - LAMBDA dataset, GPT-4o/GPT-3.5 baselines
  • Bhattacharyya et al., "Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries," WACV 2026 (arXiv 2511.20854) - frontier-model weakness on recall signals

See the analysis before you sign up

Open a complete fictional-creative dossier, then test three of your own creatives for $49 USD.

[ VIEW SAMPLE ANALYSIS ]
PIERCE THROUGH THE NOISE
Review 3 creatives — $49