Real platform data on hooks, length, and captions — tested against three scripts that actually shipped. Plus a downloadable outline template.
RevID figures verified against the source report, August 2026
73.5% of hooks fit no formula at all.
The templates are a starting point, not the finding. Most working hooks are unique.
Write the concept first. A prompt written without a concept produces a nice video that sells nothing. — the same warning from our AI thumbnails guide, applied to scripts
The thumbnails guide made this point about images: decide the single idea, the emotion, and the words before you touch a generator. Scripts have the identical failure mode, and it's more expensive — a weak script wastes the render credits, the edit, the thumbnail, and the upload slot behind it.
| What people do | What it produces |
|---|---|
| Open the generator and type a topic | A competent summary of a topic. Informative, forgettable, no reason to keep watching. |
| Write the concept first — what's the one idea, and why would someone stay | A script with a reason to exist. |
This is the distinction that separates scripts that work from scripts that are merely accurate.
An accurate script conveys the information. A working script conveys the information while giving the viewer a reason to still be there at second forty. Those are different jobs. The first is served by knowing your topic. The second is served by structure — Part 3.
Our driving-bias script answers all three before the first line: the idea is that self-assessment measures a flattering guess rather than reality; the tension is a statistic that can't be true; the takeaway is that this is arithmetic, not a character flaw.
RevID ran a classifier across 348,346 exported videos with usable script data. Here's the full distribution, not just the headline.
| Hook type | Share | RevID's example |
|---|---|---|
| Question | 12.2% | "Did you know...?" |
| Bold statement | 6.6% | Direct, opinion-led openers |
| Storytelling | 5.4% | "I used to..." |
| Statistic | 2.2% | "9 out of 10 people..." |
| Other / mixed | 73.5% | Hybrid or non-classified openings |
Source: RevID.ai, "The Anatomy of a Viral Short: What 756,851 AI-Generated Videos Tell Us About Content in 2026," March 2026. Classification by heuristic NLP classifier on the first sentence.
RevID says it directly: 73.5% of video openings don't fit neatly into the four classical hook categories, which suggests the most effective hooks are often unique, contextual, and hard to replicate with a formula.
Most script advice presents hook templates as the answer. The data says templates describe roughly a quarter of what gets made.
The useful read: question hooks are the most reliable starting point when you're stuck, not the best hook available. If you have something specific and strange to open with, use it.
The psychological mechanism is well-established: open loops. Our brains are wired to seek closure, and a well-placed question creates exactly that tension. — RevID.ai
RevID's creator takeaway is worth repeating: start with a question whenever you can, but not a rhetorical one — a genuine question your audience is already asking. The hook isn't about being clever. It's about creating a reason to keep watching.
Right now, if I asked you to rate your driving skill compared to everyone else on the road, what would you say? Above average? Way above average? — opening line, "Why You Always Think You're a Better Driver Than Everyone Else"
It opens with a direct question to the viewer, then follows immediately with the impossible statistic — "nine out of ten drivers rate themselves as better than average." RevID's example for the Statistic category is literally "9 out of 10 people..." — almost word for word the second beat of this hook.
Statistic hooks are the rarest classified category at 2.2%. This one leads with a question and lands the statistic as the seal — using the two rarest-to-formalize mechanisms together rather than picking one. The tension doesn't come from the grammar. It comes from the gap the statistic opens.
The hook doesn't stop at the question. The next lines do the actual work:
Here's the thing — statistically, that's impossible for most of us to be true. And yet, in study after study, roughly nine out of ten drivers rate themselves as better than average. Half of those people are mathematically wrong, and none of them know which half they're in.
Line one raises a claim. Line two proves the claim can't be true for most people. Line three says people believe it anyway. That's a complete open loop in three sentences, and it's why the viewer stays.
The average RevID script runs 117.5 words — roughly 60–75 seconds of spoken content. That means the hook occupies the first 10–15 words of a ~120-word story.
Ten to fifteen words. That's one sentence. For a long-form script the proportion is different, but the principle holds — your hook is a sentence, not a paragraph. Everything after it is already the body.
| Type | Use when | Watch out for |
|---|---|---|
| Question | You're stuck, or the audience is already asking it. | Rhetorical questions. "Ever wonder why?" is filler. |
| Statistic | The number is genuinely surprising or impossible. | A number that needs explaining first isn't a hook. |
| Bold statement | You have a real contrarian position you can defend. | Overclaiming, then walking it back. |
| Storytelling | The story is short and the payoff is fast. | A slow build — you don't have the runway. |
| Something else | You have a specific, strange, concrete detail. | Nothing — this is 73.5% of what works. |
| Stage | Job | Failure mode |
|---|---|---|
| Hook | Create the open loop. One sentence. | Explaining before you've earned attention. |
| Tension / mechanism | Show why the thing is true. This is the body, and it's where retention is won or lost. | Listing facts instead of building an explanation. |
| Payoff | Close the loop. Deliver what the hook promised. | Never actually answering the opening question. |
| Takeaway | Give them something to leave with. | A generic summary, or a CTA where a thought should be. |
Most weak scripts have a decent hook and a decent conclusion with a pile of facts in between.
Mechanism means explaining why, in a chain the viewer can follow. Each step should make the next one feel necessary. That's what keeps someone watching at minute four — not more information, but the sense that the explanation is going somewhere.
The structure in a real 1,187-word script, mapped against its actual section markers.
| Marker | Section | Stage |
|---|---|---|
| 0:00 | The impossible statistic | Hook |
| 1:00 | The illusion of asymmetric insight | Mechanism, part 1 |
| 2:40 | Why your brain protects your self-image | Mechanism, part 2 |
| 4:15 | The skill-vs-luck problem | The turn — reframes what came before |
| 5:45 | Why facts don't fix this | Payoff, part 1 |
| 7:00 | What actually works | Payoff, part 2 |
| 7:35 | Takeaway | Takeaway |
Worth noting because it has practical value: this script broke cleanly into seven self-contained Shorts — the statistic, asymmetric insight, self-image protection, near-miss-equals-skill, facts-don't-change-minds, what-actually-works, and the invisible-bias takeaway.
That's not a coincidence of editing. It's a property of the structure. Because each mechanism step was self-contained and had its own small payoff, each one survived being lifted out.
A script written as a continuous argument with no internal beats doesn't cut. A script written as a chain of complete steps does. That's a real argument for the four-part structure that most script guides miss: it's not just about retention, it's about downstream reuse.
| Script | Words | Hook type | Note |
|---|---|---|---|
| "Why Losing $20 Feels Worse Than Finding $20 Feels Good" | 1,101 | Scenario / contrast — two mirrored images (finding vs. losing the same $20) before any question or statistic lands | Two-concept structure: loss aversion, then the endowment effect as a related second mechanism, rather than one idea explored linearly |
| "Why Online Arguments Escalate Faster Than In-Person Ones" | 1,038 | Contrast — a calm dinner-table disagreement set against a comment-section blowup, before either is named | Four parallel mechanisms (disinhibition, dehumanization, audience effect, rumination) rather than one deepening explanation — proof the four-part shape holds even with a wider, list-like mechanism section |
All three scripts hit different hook mechanics and different mechanism shapes — one linear causal chain, one two-concept pairing, one four-item parallel structure — but all three keep the same hook → mechanism → turn → payoff → takeaway skeleton. That's the actual claim of Part 3: the shape is reusable across very different content, not just repeatable within one style.
| Duration | Share of videos | RevID's context |
|---|---|---|
| 0–15 seconds | 4.4% | Ultra-short hooks, teasers, memes |
| 16–30 seconds | 11.9% | Classic short-form range |
| 31–60 seconds | 31.1% | Sweet spot for educational content |
| 61–90 seconds | 16.9% | Extended storytelling |
| 90+ seconds | 35.7% | Long-form shorts — podcasts, explainers |
Source: RevID.ai 2026 data report. Average exported video runs 115.5 seconds.
The era of the 15-second clip as the default format is over. — RevID.ai
90+ seconds at 35.7% and 31–60 seconds at 31.1% together account for over two-thirds of all exported content. The sub-30-second range — what most people picture when they hear "short video" — is 16.3%. A minority.
RevID's explanation is that platforms now reward watch time alongside completion rate, and creators responded.
Don't artificially cut your content to fit an outdated template. If your story needs 90 seconds, give it 90 seconds. — RevID.ai creator takeaway
This is where most script advice does damage. The instruction to "keep it under 30 seconds" is a template inherited from a version of the platform that no longer exists, and following it forces writers to strip out the mechanism — the exact part that holds attention.
We tested this directly: an 880-word script rendered at 352 seconds in RevID. That's exactly 150 words per minute — the number we've used for every script since.
For reference, RevID's platform-wide average implies something slower (117.5 words for 60–75 seconds is roughly 94–117 wpm). Voice choice and pacing settings affect this — your own measured number describes your renders, not the platform average.
Words for target runtime = runtime (seconds) ÷ 60 × 150
| Target runtime | Words at 150 wpm |
|---|---|
| 30 seconds | 75 |
| 60 seconds | 150 |
| 90 seconds | 225 |
| 3 minutes | 450 |
| 5 minutes | 750 |
| 8 minutes | 1,200 |
| 10 minutes | 1,500 |
Write the script, render it once, then measure the actual output against your word count — don't assume a runtime and schedule around it. A single test render is what gave us the 150 wpm figure above, and it's held consistently across every script since.
| YouTube long-form | YouTube Shorts | TikTok | |
|---|---|---|---|
| Ideal length | 8+ min for mid-roll eligibility | Up to 3 min counts as a Short | Longer favored — Creator Rewards requires 1+ min |
| Hook window | First 15–30 seconds | First 3 seconds | First 3 seconds |
| Aspect ratio | 16:9 | 9:16 | 9:16 |
| Caption style | Optional but recommended | Animated, word-by-word | Animated, word-by-word |
| Script density | Can breathe — build the mechanism | Compressed — one beat only | Compressed — one beat only |
| Structure | Full four-part shape | Hook + payoff. Mechanism is usually one line. | Same |
| Payoff timing | Can be delayed and earned | Must arrive fast | Must arrive fast |
| CTA norms | Accepted | Minimal — the format punishes it | Minimal |
| Discovery | Search and suggested | Feed | Feed |
| What it's for | Revenue and depth | Reach and discovery | Reach and discovery |
Compiled from the RevID 2026 report (9:16 at 77.9% of exports; caption norms), and platform guidance in our Shorts monetization guide and VidIQ guide.
A long-form script builds a mechanism over several steps and earns its payoff. Compressing that into 45 seconds produces a script that gestures at an explanation without delivering one — all setup, no substance.
A Short should be one beat: a hook and its payoff, with the mechanism as a single line rather than a section. Which is exactly what our seven-beat cut did — each Short took one complete step from the long-form and let it stand alone. That's the correct pattern, not summarizing the whole video into a minute.
RevID's aspect ratio data confirms the default: 9:16 vertical at 77.9%, 16:9 landscape at 20.9%, square at 1.1%. If you're not shooting vertical-first, you're working against the algorithm on every major platform. For a script, that means writing the Shorts version knowing the frame is tall and captions occupy real estate.
| Metric | Value |
|---|---|
| Videos with captions enabled | 86.6% |
| Of captioned videos: animated | 96.8% |
| Of captioned videos: static | 3.2% |
| Social video watched without sound | 85% — the long-cited figure RevID opens with |
Source: RevID.ai 2026 data report.
Animated captions do two things simultaneously: they make content accessible for silent viewers, and they create a visual rhythm that actively guides attention. They're not a caption style. They're a retention mechanism. — RevID.ai
Captions get framed as an accessibility feature that also happens to help. The data suggests the causality runs the other way for most creators — 96.8% animated adoption is not an accessibility decision, it's a retention decision.
Word-by-word reveal creates a visual pulse synchronized to the narration. The eye tracks it. That's attention capture, not just readability. Both reasons are real, and accessibility deserves its due — but the reason 96.8% chose animated over static is not accessibility. Static captions are equally accessible.
RevID notes caption adoption was near-universal in H1 2025 before settling around 84–87% from August onward, and suggests the softening may reflect an influx of musicians and entertainment creators who prioritize visual aesthetics over accessibility.
So the 86.6% is a platform average across content types. For explainer content specifically — our niche — the case for captions is stronger than the average suggests, not weaker.
Generation is a first draft, not a final one. — our faceless-channel guide, on AI output
Our thumbnails guide said not to publish the first image. The faceless-channel guide said not to publish the first render. The script is where that principle matters most, because everything downstream is built on it.
It's also the stage where the July 2025 inauthentic content policy bites. YouTube's named example is channels that upload narrative stories with only superficial differences between them — that's a script-level problem, not a production-level one.
| Step | What happens |
|---|---|
| 1 | Research the topic yourself. Pull the actual mechanism or finding from a real source. |
| 2 | Write a short outline — not prose. The accurate skeleton: hook, mechanism, real-world example, takeaway. |
| 3 | Feed that outline to RevID to expand into a full script in your format and tone. |
| 4 | Read the generated script once before rendering, checking specifically for anything added or reworded that isn't in your outline. |
The check isn't "does this read well." It's: what's in here that I didn't put in my outline?
For a research-based niche, that's a factual-accuracy check — a misattributed study or an overstated finding is exactly what an informed audience catches. For any niche, it's an originality check. It's also fast, since you're comparing against your own outline, not re-researching.
| Cut | Why |
|---|---|
| Throat-clearing before the hook | "In today's video we're going to talk about..." spends your only guaranteed attention. |
| Restatements | AI drafts restate the thesis at every section boundary. Keep one. |
| Hedging stacks | "It could be argued that in some cases it may be possible that..." Pick a claim. |
| Facts that don't serve the mechanism | Interesting isn't the bar. Does it move the explanation forward? |
| Adjective pileups | AI writes "fascinating, surprising, remarkable." Let the content be those things. |
| Generic transitions | "Now let's dive into..." Cut it and start the next sentence. |
| Anything you can't source | Especially in a research-based niche. |
| The CTA, sometimes | Our driving-bias script's body ends with no CTA and is stronger for it. |
Read it out loud. Every place you stumble is a place the narration will sound wrong, and every place you get bored is a place the viewer will leave.
This catches things silent reading doesn't: clauses that are too long, repetition you skimmed past, and sentences that parse on the page but not in the ear. For AI-generated drafts it's especially useful, because generated prose tends to read smoothly and speak awkwardly.
The deliverable for this page. Word-count guardrails are set at 150 wpm — swap in your own measured pace once you have one.
Google Docs-ready .docx — hook/mechanism/turn/payoff/takeaway, with the Shorts variant included.
| Stage | Guardrail at ~1,200 words | What the script actually did |
|---|---|---|
| Hook | 10–15 words + 2 sealing lines | Question, then the impossible statistic — three sentences |
| Mechanism 1 | ~240 words | Asymmetric insight, named and evidenced |
| Mechanism 2 | ~300 words | Self-image protection, three overlapping reasons |
| The turn | ~240 words | Skill-vs-luck — near-misses reframed as luck, not competence |
| Payoff | ~240 words | Why facts don't fix it, then what actually works |
| Takeaway | ~120 words | Not a character flaw — just the math of self-perception |
The same shape works at 60 seconds and 10 minutes. Fixed counts would only serve one length, and Part 4 is explicitly about not forcing a template length.
At 150 words (60 seconds), the mechanism steps collapse into one line each and the turn may disappear. That's fine — that's a Short. At 1,200+ words the full shape has room to breathe. Same template, different scale.
Same tool for every case study documented here. 20% off with code THEAIBUILDLAB.
Try RevID.ai →| Mistake | Consequence |
|---|---|
| Prompting before you have a concept | A competent summary nobody stays for. |
| Treating hook templates as the answer | 73.5% of hooks fit no formula. Templates are a fallback, not a target. |
| Writing a rhetorical question as a hook | "Ever wonder why?" is filler, not an open loop. |
| A hook longer than a sentence | RevID's math: 10–15 words of a ~120-word script. |
| Underwriting the mechanism | Facts in a pile don't retain. A causal chain does. |
| No turn | The script flattens out. The turn is the retention beat. |
| Never closing the loop | If the hook asked something, answer it. |
| A summary instead of a takeaway | Reframe, don't recap. |
| Cutting to fit a template length | 90+ seconds is the largest bucket at 35.7%. Sub-30 is 16.3%. |
| Planning a calendar on an estimated runtime | Measure your own render once — words-per-minute varies by voice and settings. |
| Compressing a long-form script into a Short | A Short is one beat, not a summary. |
| Static captions, or none | 86.6% caption, and 96.8% of those animate. It's a retention mechanism. |
| Long clauses in a captioned script | Word-by-word reveal turns them into a crawl. |
| Publishing the AI's first draft | Same principle as the first render. The script is where originality lives. |
| Not diffing against your outline | That's where added inaccuracies hide. |
| Never reading it aloud | Free, and it catches what silent reading misses. |
The RevID data describes RevID's own creator base specifically — genuine first-party data with a disclosed methodology, but not independently audited, and the report itself states it may not generalize to all short-form video. Hook classification is heuristic, not human-reviewed, so the 73.5% "other/mixed" category likely contains both genuinely novel hooks and classifier misses — we can't distinguish those from the published data. The 150 wpm pacing figure is measured from our own renders, not RevID's platform average, and will vary by voice, style, and settings. This is not a guarantee that following this structure produces a specific outcome — RevID's own FAQ states no tool can guarantee views, and neither can a structure.