DORU-1:preview
A multi-modal creative evaluation engine for pre-execution ad analysis. Scores video and static creatives before a dollar of media spend is committed.
> 01. Abstract
DORU-1:preview is a multi-modal creative evaluation engine that analyzes static and video advertisements before media capital is deployed. A single asset produces a structured dossier of diagnostic panels: visual saliency and attention curves, memorability scores, transcripts, emotion fingerprints, hook grades, copy and CTA analysis, weak-moment flags, advisory compliance screening, and deterministic audit receipts.
The core contribution is the memorability head: a compact neural predictor trained on human long-term recall data (LAMBDA, WACV 2025; 1,749 participants, 2,205 advertisements). On the LAMBDA test split it achieves Spearman ρ = 0.413 (static) and ρ = 0.345 (video) - 54% and 45% of the human split-half consistency ceiling (ρ = 0.77).
We additionally report the first published benchmark of 2026 frontier models on ad memorability: Gemini 3.6 Flash (0.268), GPT-5.6 Luna (0.247), Grok 4.5 (0.217), Sonnet 5 (0.144), and Gemini 3.1 Pro (0.131) all score below both DORU-1 heads on the same 219-ad protocol.
> 02. The Benchmark
DORU-1 does not predict engagement, click-through, or conversion. It predicts and measures the components that precede them: attention, memorability, emotional response, message clarity, and pacing. Activation is not attention. Attention is not engagement. Engagement is not purchase intent. Each layer is a separate problem, and conflating them is how creative-scoring products fail.
See how these measurements become ranked creative decisions in the public sample dossier.
LAMBDA test split - Spearman rank correlation
| Model | ρ | n | 95% CI |
|---|---|---|---|
| Human split-half consistency (ceiling) | 0.77 | - | - |
| DORU-1 static | 0.413 | 214 | [0.29, 0.52] |
| DORU-1 video | 0.345 | 214 | [0.22, 0.46] |
| Gemini 3.6 Flash (thinking) | 0.268 | 219 | [0.140, 0.386] |
| GPT-5.6 Luna (high) | 0.247 | 219 | [0.118, 0.367] |
| Grok 4.5 (fast) | 0.217 | 219 | [0.087, 0.340] |
| Sonnet 5 (medium) | 0.144 | 219 | [0.011, 0.271] |
| Gemini 3.1 Pro | 0.131 | 219 | [−0.001, 0.259] |
| GPT-4o 10-shot (published) | 0.18 | 219 | - |
| GPT-3.5 10-shot (published) | 0.06 | 219 | - |
Ratio to the human consistency ceiling
The frontier plateau
Three generations of generalist models: 0.06 (GPT-3.5, 2023) to 0.18 (GPT-4o, 2025) to a 0.13–0.27 band across the 2026 frontier. The gap to the human ceiling has not moved. A dedicated head, trained on the right data, did.
Reading the table
- All five 2026 frontier models score below both DORU-1 heads in point estimate, with DORU-1 using a compact dedicated regressor — orders of magnitude fewer parameters — against frontier models with ~100B+ active parameters.
- Pairwise gaps vs DORU-1: the differences vs Grok 4.5, Sonnet 5, and Gemini 3.1 Pro are statistically robust; the gaps vs Gemini 3.6 Flash and GPT-5.6 Luna are directionally consistent but individually within sampling noise at n≈214. Larger test sets, a planned scale-up path, will tighten these intervals.
- Why: memorability is an unintuitive task. As the TOT2MEM study (Bhattacharyya et al., WACV 2026) found, general multimodal reasoning and world knowledge may not be sufficient for predicting recall signals. A dedicated head trained on the right labels extracts signal generalists cannot.
- Category position: DORU-1 is the only full-stack creative evaluation platform benchmarked on this protocol; no comparable multi-panel system exists to compare against. The comparison above is against general-purpose frontier models and the human consistency ceiling.
- Protocol notes: Sonnet 5 compressed its output scale (15–65, 22 unique values), so its ρ is likely an underestimate. Gemini 3.1 Pro's result is borderline (p=0.052). All runs are single-shot; the literature baselines averaged 3 seeds. Grok 4.5's initial run fabricated sequential IDs and was discarded; the reported value is a clean re-run.
> 03. System Overview
DORU-1 is served within the Spear platform. Creatives are uploaded and analyzed asynchronously, and results are returned as structured dossiers that drive the buyer/publisher review exchange. All evaluation compute runs on Spear Systems infrastructure.
Analysis pipeline
Analysis panels (v1)
| Panel | Output |
|---|---|
| Video saliency | 60-frame attention curve + heatmap |
| Static saliency | 256-cell attention heatmap |
| Memorability (video) | score 0–1 + metrics |
| Memorability (static) | score 0–1 + metrics, text-augmented |
| Transcript | word/segment timing |
| Emotion fingerprint | 4D vector, confidence-tagged |
| ABCD features | hook, pacing, CTA, faces, speech, supers |
| Hook grade | PASS/FAIL |
| Copy/CTA analysis | headline/body/CTA assessment |
| Weak-moment flags | timestamped attention-dip flags |
| Compliance screening | advisory flags, timecoded |
| Audit receipt | deterministic, replay-verifiable |
The language analysis layer uses a fallback chain of hosted and self-hosted analysis models. Operational guardrails bound queue depth and per-partner throughput.
> 04. Memorability Heads
Architecture
The memorability heads are compact neural regressors layered on frozen, pre-trained visual backbones, trained on human long-term recall data. The video head fuses raw low-level features with scene-level and event-level representations; the static head is augmented with a language signal derived from the creative's copy, transcript, and on-screen text.
Training
Labels come from LAMBDA (Si et al., WACV 2025): long-term brand-recall scores from 1,749 participants across 2,205 real advertisements (276 brands, 113 industries, avg 33s). Split: 1,964 train / 219 test.
Trained with a standard regression objective on the train split, with a held-out validation split for checkpoint selection and raw-feature normalization fit on train only. Backbones remain frozen; only the compact regressor trains.
Results
| Head | Test Spearman (n=214) | % of ceiling (0.77) |
|---|---|---|
| Static (text-augmented) | 0.413 | 54% |
| Video | 0.345 | 45% |
> 05. Saliency and Attention
DORU-1 maps viewer attention second-by-second. Video assets produce a 60-frame attention curve and heatmap; static assets produce a 256-cell attention heatmap.
The attention curve is converted to a saliency-entropy trace; sustained entropy dips mark moments where visual interest collapses. These are surfaced as timestamped weak-moment flags: structural signals for editors, telling them where to insert a pattern interrupt or re-hook. The model identifies where; the creative team decides what.
The first-3-second window is evaluated against a rubric (face presence, bold supers, clear problem statement, auditory alignment) yielding a PASS/FAIL hook grade. This is a structural check, not a prediction of retention.
> 06. Language Analysis Layer
The language analysis layer produces the qualitative panels:
- Emotion fingerprint: a 4D vector (trust, excitement, discomfort, memorability) with per-dimension confidence tags. Subcortical reward signals are explicitly confidence-tagged: they are proxies, not measurements.
- ABCD creative features: hook, pacing, CTA, faces, speech, supers, inferred from frames and audio.
- Copy/CTA analysis: headline, body, and call-to-action assessment against the creative's own claims.
The layer is decoupled from the memorability heads: swapping the analysis model does not change the memorability score, and vice versa.
> 07. Compliance Screening (Advisory)
DORU-1 scans transcripts and on-screen text for claim patterns (financial, medical, outcome-based) and surfaces timecoded flags into the review exchange.
The screening is advisory by design. A flagged asset enters the review queue, the publisher acknowledges with a comment, and the buyer retains final compliance authority. It accelerates the compliance workflow; it does not replace standard compliance checks.
> 08. Audit Receipts
Every analysis produces a deterministic, replay-verifiable receipt binding the asset, the model version, and the results. Any change to the analysis configuration invalidates prior receipts by construction. Scores are reproducible or they are not published.
CREATIVE ID : 8f3a-2c91-4e77-9b10
ENGINE : doru-1:preview
MODEL VERSION : v0.1.0
MEDIA HASH : a1b2c3...f0e1
PANELS : saliency · memorability · emotion
· transcript · hook · compliance
RECEIPT : 9f86d081...0f00a08
VERIFICATION : REPLAY VERIFIED - inputs and outputs
bound; any config change invalidates
STATUS : APPROVED WITH CONDITIONS> 09. Syntagma-1 - Roadmap
Syntagma-1:preview is the planned enterprise engine: a neural-response layer and a macro-context layer, gated on infrastructure readiness. Architecture details are available under NDA.
Neither layer is live in DORU-1. DORU-1 is explicitly creative-only; macro context is excluded by design.
| Capability | DORU-1:preview | Syntagma-1 |
|---|---|---|
| Status | LIVE | IN DEVELOPMENT |
| Input modalities | Static + video | Full multi-modal |
| Memorability scoring | ✓ | ✓ |
| Saliency / attention | ✓ | ✓ |
| Compliance screening | advisory | advisory |
| Neural response layer | - | planned |
| Macro context layer | excluded by design | planned |
| Primary use case | High volume | High capital |
> 10. Risk Factors
- 01Scores are creative-only. DORU-1 evaluates the creative itself; seasonality, market, and broader context are not yet part of scoring.
- 02Compliance screening is advisory. The buyer retains final compliance authority; screening accelerates review but does not gate.
- 03Norm benchmarks are still building. Percentile baselines strengthen as more creatives run through the system.
> 11. Licensing and Corpus Strategy
All model components are commercially licensed for use, verified at the weights level, not just the repository license. The benchmark dataset (LAMBDA) is publicly released under a permissive license (labels MIT-licensed; source footage remains under the original creators' licenses). No non-commercial or research-only weights are used in the shipped engine.
Partner data usage is governed by platform terms agreed at onboarding.
> 12. Conclusion
DORU-1:preview is a pre-execution creative evaluation engine with real, measured, reproducible results: memorability heads at 54% and 45% of the human consistency ceiling, and the first published benchmark showing 2026 frontier models scoring below both heads on the same protocol.
It is not the finished product. Norm baselines await corpus volume, and the macro layer is roadmap. The architecture, the measurements, and the limitations are documented so the claims can be checked, not just believed.
> 13. References
- Si et al., "Long-Term Ad Memorability: Understanding & Generating Memorable Ads," WACV 2025 (arXiv 2309.00378) - LAMBDA dataset, GPT-4o/GPT-3.5 baselines
- Bhattacharyya et al., "Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries," WACV 2026 (arXiv 2511.20854) - frontier-model weakness on recall signals
See the analysis before you sign up
Open a complete fictional-creative dossier, then test three of your own creatives for $49 USD.
[ VIEW SAMPLE ANALYSIS ]