Guide · Built on 756,851 videos and 3 real scripts

How to Write Video Scripts That Actually Work

Real platform data on hooks, length, and captions — tested against three scripts that actually shipped. Plus a downloadable outline template.

RevID figures verified against the source report, August 2026

73.5% of hooks fit no formula at all.

The templates are a starting point, not the finding. Most working hooks are unique.

Affiliate disclosure Some links on this page, including to RevID.ai, are affiliate links. That doesn't change what we write — the data below comes from RevID's own published report, and the script analysis is our own real, published work.
Part 1

Why most scripts fail before they're written

Write the concept first. A prompt written without a concept produces a nice video that sells nothing. — the same warning from our AI thumbnails guide, applied to scripts

The thumbnails guide made this point about images: decide the single idea, the emotion, and the words before you touch a generator. Scripts have the identical failure mode, and it's more expensive — a weak script wastes the render credits, the edit, the thumbnail, and the upload slot behind it.

What people doWhat it produces
Open the generator and type a topicA competent summary of a topic. Informative, forgettable, no reason to keep watching.
Write the concept first — what's the one idea, and why would someone stayA script with a reason to exist.
🎯 A script is not an information delivery document. It's a retention document.

This is the distinction that separates scripts that work from scripts that are merely accurate.

An accurate script conveys the information. A working script conveys the information while giving the viewer a reason to still be there at second forty. Those are different jobs. The first is served by knowing your topic. The second is served by structure — Part 3.

Three questions to answer before writing a word

  1. What is the one idea? If you can't say it in a sentence, the script will wander.
  2. Why would someone stay? What tension, question, or surprise carries them past the first ten seconds?
  3. What do they leave with? If the answer is "they learned about a topic," that's not a takeaway. A takeaway is something they can use, repeat, or feel.

Our driving-bias script answers all three before the first line: the idea is that self-assessment measures a flattering guess rather than reality; the tension is a statistic that can't be true; the takeaway is that this is arithmetic, not a character flaw.

Part 2

The hook, backed by real data

What RevID found

RevID ran a classifier across 348,346 exported videos with usable script data. Here's the full distribution, not just the headline.

Hook typeShareRevID's example
Question12.2%"Did you know...?"
Bold statement6.6%Direct, opinion-led openers
Storytelling5.4%"I used to..."
Statistic2.2%"9 out of 10 people..."
Other / mixed73.5%Hybrid or non-classified openings

Source: RevID.ai, "The Anatomy of a Viral Short: What 756,851 AI-Generated Videos Tell Us About Content in 2026," March 2026. Classification by heuristic NLP classifier on the first sentence.

🚩 Don't skip the 73.5% — it's the most honest number in the report

RevID says it directly: 73.5% of video openings don't fit neatly into the four classical hook categories, which suggests the most effective hooks are often unique, contextual, and hard to replicate with a formula.

Most script advice presents hook templates as the answer. The data says templates describe roughly a quarter of what gets made.

The useful read: question hooks are the most reliable starting point when you're stuck, not the best hook available. If you have something specific and strange to open with, use it.

Why questions work when they work

The psychological mechanism is well-established: open loops. Our brains are wired to seek closure, and a well-placed question creates exactly that tension. — RevID.ai

RevID's creator takeaway is worth repeating: start with a question whenever you can, but not a rhetorical one — a genuine question your audience is already asking. The hook isn't about being clever. It's about creating a reason to keep watching.

Our own hook, and what it actually is

Right now, if I asked you to rate your driving skill compared to everyone else on the road, what would you say? Above average? Way above average? — opening line, "Why You Always Think You're a Better Driver Than Everyone Else"
🎯 Notice something: that's closer to a question hook than a statistic hook

It opens with a direct question to the viewer, then follows immediately with the impossible statistic — "nine out of ten drivers rate themselves as better than average." RevID's example for the Statistic category is literally "9 out of 10 people..." — almost word for word the second beat of this hook.

Statistic hooks are the rarest classified category at 2.2%. This one leads with a question and lands the statistic as the seal — using the two rarest-to-formalize mechanisms together rather than picking one. The tension doesn't come from the grammar. It comes from the gap the statistic opens.

The structural detail underneath it

The hook doesn't stop at the question. The next lines do the actual work:

Here's the thing — statistically, that's impossible for most of us to be true. And yet, in study after study, roughly nine out of ten drivers rate themselves as better than average. Half of those people are mathematically wrong, and none of them know which half they're in.

Line one raises a claim. Line two proves the claim can't be true for most people. Line three says people believe it anyway. That's a complete open loop in three sentences, and it's why the viewer stays.

The hook budget

⚠️ RevID's hook math is a useful constraint

The average RevID script runs 117.5 words — roughly 60–75 seconds of spoken content. That means the hook occupies the first 10–15 words of a ~120-word story.

Ten to fifteen words. That's one sentence. For a long-form script the proportion is different, but the principle holds — your hook is a sentence, not a paragraph. Everything after it is already the body.

Hook types, with when to use each

TypeUse whenWatch out for
QuestionYou're stuck, or the audience is already asking it.Rhetorical questions. "Ever wonder why?" is filler.
StatisticThe number is genuinely surprising or impossible.A number that needs explaining first isn't a hook.
Bold statementYou have a real contrarian position you can defend.Overclaiming, then walking it back.
StorytellingThe story is short and the payoff is fast.A slow build — you don't have the runway.
Something elseYou have a specific, strange, concrete detail.Nothing — this is 73.5% of what works.
Part 3

Structure that holds attention

The four-part shape

StageJobFailure mode
HookCreate the open loop. One sentence.Explaining before you've earned attention.
Tension / mechanismShow why the thing is true. This is the body, and it's where retention is won or lost.Listing facts instead of building an explanation.
PayoffClose the loop. Deliver what the hook promised.Never actually answering the opening question.
TakeawayGive them something to leave with.A generic summary, or a CTA where a thought should be.
🎯 The mechanism section is the one people underwrite

Most weak scripts have a decent hook and a decent conclusion with a pile of facts in between.

Mechanism means explaining why, in a chain the viewer can follow. Each step should make the next one feel necessary. That's what keeps someone watching at minute four — not more information, but the sense that the explanation is going somewhere.

Worked example — the driving-bias script

The structure in a real 1,187-word script, mapped against its actual section markers.

MarkerSectionStage
0:00The impossible statisticHook
1:00The illusion of asymmetric insightMechanism, part 1
2:40Why your brain protects your self-imageMechanism, part 2
4:15The skill-vs-luck problemThe turn — reframes what came before
5:45Why facts don't fix thisPayoff, part 1
7:00What actually worksPayoff, part 2
7:35TakeawayTakeaway

What the script does well, specifically

The seven-beat property

Worth noting because it has practical value: this script broke cleanly into seven self-contained Shorts — the statistic, asymmetric insight, self-image protection, near-miss-equals-skill, facts-don't-change-minds, what-actually-works, and the invisible-bias takeaway.

🎯 A well-structured long-form script is a Shorts pipeline

That's not a coincidence of editing. It's a property of the structure. Because each mechanism step was self-contained and had its own small payoff, each one survived being lifted out.

A script written as a continuous argument with no internal beats doesn't cut. A script written as a chain of complete steps does. That's a real argument for the four-part structure that most script guides miss: it's not just about retention, it's about downstream reuse.

Two more real examples

ScriptWordsHook typeNote
"Why Losing $20 Feels Worse Than Finding $20 Feels Good"1,101Scenario / contrast — two mirrored images (finding vs. losing the same $20) before any question or statistic landsTwo-concept structure: loss aversion, then the endowment effect as a related second mechanism, rather than one idea explored linearly
"Why Online Arguments Escalate Faster Than In-Person Ones"1,038Contrast — a calm dinner-table disagreement set against a comment-section blowup, before either is namedFour parallel mechanisms (disinhibition, dehumanization, audience effect, rumination) rather than one deepening explanation — proof the four-part shape holds even with a wider, list-like mechanism section

All three scripts hit different hook mechanics and different mechanism shapes — one linear causal chain, one two-concept pairing, one four-item parallel structure — but all three keep the same hook → mechanism → turn → payoff → takeaway skeleton. That's the actual claim of Part 3: the shape is reusable across very different content, not just repeatable within one style.

Part 4

Length and pacing

The length data

DurationShare of videosRevID's context
0–15 seconds4.4%Ultra-short hooks, teasers, memes
16–30 seconds11.9%Classic short-form range
31–60 seconds31.1%Sweet spot for educational content
61–90 seconds16.9%Extended storytelling
90+ seconds35.7%Long-form shorts — podcasts, explainers

Source: RevID.ai 2026 data report. Average exported video runs 115.5 seconds.

The era of the 15-second clip as the default format is over. — RevID.ai
🎯 Two-thirds of content is over 31 seconds, and the biggest single bucket is 90+

90+ seconds at 35.7% and 31–60 seconds at 31.1% together account for over two-thirds of all exported content. The sub-30-second range — what most people picture when they hear "short video" — is 16.3%. A minority.

RevID's explanation is that platforms now reward watch time alongside completion rate, and creators responded.

The lesson that matters more than the numbers

Don't artificially cut your content to fit an outdated template. If your story needs 90 seconds, give it 90 seconds. — RevID.ai creator takeaway

This is where most script advice does damage. The instruction to "keep it under 30 seconds" is a template inherited from a version of the platform that no longer exists, and following it forces writers to strip out the mechanism — the exact part that holds attention.

Our own pacing math

🚩 Use your measured number, not a rule of thumb

We tested this directly: an 880-word script rendered at 352 seconds in RevID. That's exactly 150 words per minute — the number we've used for every script since.

For reference, RevID's platform-wide average implies something slower (117.5 words for 60–75 seconds is roughly 94–117 wpm). Voice choice and pacing settings affect this — your own measured number describes your renders, not the platform average.

The practical table

Words for target runtime = runtime (seconds) ÷ 60 × 150

Target runtimeWords at 150 wpm
30 seconds75
60 seconds150
90 seconds225
3 minutes450
5 minutes750
8 minutes1,200
10 minutes1,500
✅ Measure before you plan a calendar

Write the script, render it once, then measure the actual output against your word count — don't assume a runtime and schedule around it. A single test render is what gave us the 150 wpm figure above, and it's held consistently across every script since.

Part 5

Platform differences

YouTube long-formYouTube ShortsTikTok
Ideal length8+ min for mid-roll eligibilityUp to 3 min counts as a ShortLonger favored — Creator Rewards requires 1+ min
Hook windowFirst 15–30 secondsFirst 3 secondsFirst 3 seconds
Aspect ratio16:99:169:16
Caption styleOptional but recommendedAnimated, word-by-wordAnimated, word-by-word
Script densityCan breathe — build the mechanismCompressed — one beat onlyCompressed — one beat only
StructureFull four-part shapeHook + payoff. Mechanism is usually one line.Same
Payoff timingCan be delayed and earnedMust arrive fastMust arrive fast
CTA normsAcceptedMinimal — the format punishes itMinimal
DiscoverySearch and suggestedFeedFeed
What it's forRevenue and depthReach and discoveryReach and discovery

Compiled from the RevID 2026 report (9:16 at 77.9% of exports; caption norms), and platform guidance in our Shorts monetization guide and VidIQ guide.

🎯 A Short is not a shortened long-form script

A long-form script builds a mechanism over several steps and earns its payoff. Compressing that into 45 seconds produces a script that gestures at an explanation without delivering one — all setup, no substance.

A Short should be one beat: a hook and its payoff, with the mechanism as a single line rather than a section. Which is exactly what our seven-beat cut did — each Short took one complete step from the long-form and let it stand alone. That's the correct pattern, not summarizing the whole video into a minute.

RevID's aspect ratio data confirms the default: 9:16 vertical at 77.9%, 16:9 landscape at 20.9%, square at 1.1%. If you're not shooting vertical-first, you're working against the algorithm on every major platform. For a script, that means writing the Shorts version knowing the frame is tall and captions occupy real estate.

Part 6

Captions as a retention mechanism

MetricValue
Videos with captions enabled86.6%
Of captioned videos: animated96.8%
Of captioned videos: static3.2%
Social video watched without sound85% — the long-cited figure RevID opens with

Source: RevID.ai 2026 data report.

Animated captions do two things simultaneously: they make content accessible for silent viewers, and they create a visual rhythm that actively guides attention. They're not a caption style. They're a retention mechanism. — RevID.ai
🎯 That last line is the whole section, and it's RevID's own wording

Captions get framed as an accessibility feature that also happens to help. The data suggests the causality runs the other way for most creators — 96.8% animated adoption is not an accessibility decision, it's a retention decision.

Word-by-word reveal creates a visual pulse synchronized to the narration. The eye tracks it. That's attention capture, not just readability. Both reasons are real, and accessibility deserves its due — but the reason 96.8% chose animated over static is not accessibility. Static captions are equally accessible.

What this means when you're writing

⚠️ One honest caveat on the caption data

RevID notes caption adoption was near-universal in H1 2025 before settling around 84–87% from August onward, and suggests the softening may reflect an influx of musicians and entertainment creators who prioritize visual aesthetics over accessibility.

So the 86.6% is a platform average across content types. For explainer content specifically — our niche — the case for captions is stronger than the average suggests, not weaker.

Part 7

Editing your own script, or the AI's first draft

Generation is a first draft, not a final one. — our faceless-channel guide, on AI output
🚩 The same warning as "don't publish the first AI render," applied one stage earlier

Our thumbnails guide said not to publish the first image. The faceless-channel guide said not to publish the first render. The script is where that principle matters most, because everything downstream is built on it.

It's also the stage where the July 2025 inauthentic content policy bites. YouTube's named example is channels that upload narrative stories with only superficial differences between them — that's a script-level problem, not a production-level one.

Our own workflow

StepWhat happens
1Research the topic yourself. Pull the actual mechanism or finding from a real source.
2Write a short outline — not prose. The accurate skeleton: hook, mechanism, real-world example, takeaway.
3Feed that outline to RevID to expand into a full script in your format and tone.
4Read the generated script once before rendering, checking specifically for anything added or reworded that isn't in your outline.
🎯 Step 4 is the whole safeguard

The check isn't "does this read well." It's: what's in here that I didn't put in my outline?

For a research-based niche, that's a factual-accuracy check — a misattributed study or an overstated finding is exactly what an informed audience catches. For any niche, it's an originality check. It's also fast, since you're comparing against your own outline, not re-researching.

What to cut

CutWhy
Throat-clearing before the hook"In today's video we're going to talk about..." spends your only guaranteed attention.
RestatementsAI drafts restate the thesis at every section boundary. Keep one.
Hedging stacks"It could be argued that in some cases it may be possible that..." Pick a claim.
Facts that don't serve the mechanismInteresting isn't the bar. Does it move the explanation forward?
Adjective pileupsAI writes "fascinating, surprising, remarkable." Let the content be those things.
Generic transitions"Now let's dive into..." Cut it and start the next sentence.
Anything you can't sourceEspecially in a research-based niche.
The CTA, sometimesOur driving-bias script's body ends with no CTA and is stronger for it.

What to keep and protect

✅ The single most effective script edit costs nothing

Read it out loud. Every place you stumble is a place the narration will sound wrong, and every place you get bored is a place the viewer will leave.

This catches things silent reading doesn't: clauses that are too long, repetition you skimmed past, and sentences that parse on the page but not in the ear. For AI-generated drafts it's especially useful, because generated prose tends to read smoothly and speak awkwardly.

Part 8

The template

The deliverable for this page. Word-count guardrails are set at 150 wpm — swap in your own measured pace once you have one.

Video Script Outline Template

Google Docs-ready .docx — hook/mechanism/turn/payoff/takeaway, with the Shorts variant included.

Download .docx →
SCRIPT: [working title]
Target runtime: ___   Target words: runtime(sec) ÷ 60 × 150
Platform: [ long-form / Shorts / TikTok ]

THE ONE IDEA (one sentence, before you write anything):
_________________________________

WHY WOULD SOMEONE STAY?
_________________________________

WHAT DO THEY LEAVE WITH?
_________________________________



[HOOK] 10–15 words — one sentence
Open loop. Question, statistic, or something specific.
Follow with 1–2 lines that seal the loop.

[MECHANISM — STEP 1] ~20% of total words
Name the thing. Evidence it.

[MECHANISM — STEP 2] ~25% of total words
The causal chain. Why does this happen?
Each reason different in kind, not a list.

[THE TURN] ~20% of total words
Complicate it. The thing the viewer did not expect.
This is the retention beat.

[PAYOFF] ~20% of total words
Close the loop opened in the hook.
Bonus: implicate the viewer.

[TAKEAWAY] ~10% of total words
Reframe, do not summarize.
CTA optional — consider none.



BEAT CHECK: can each mechanism step stand alone as a Short?
READ-ALOUD CHECK: done? ___
OUTLINE DIFF: anything here I did not put in my outline? ___

Worked against the driving-bias script

StageGuardrail at ~1,200 wordsWhat the script actually did
Hook10–15 words + 2 sealing linesQuestion, then the impossible statistic — three sentences
Mechanism 1~240 wordsAsymmetric insight, named and evidenced
Mechanism 2~300 wordsSelf-image protection, three overlapping reasons
The turn~240 wordsSkill-vs-luck — near-misses reframed as luck, not competence
Payoff~240 wordsWhy facts don't fix it, then what actually works
Takeaway~120 wordsNot a character flaw — just the math of self-perception
🎯 Why percentages, not fixed word counts

The same shape works at 60 seconds and 10 minutes. Fixed counts would only serve one length, and Part 4 is explicitly about not forcing a template length.

At 150 words (60 seconds), the mechanism steps collapse into one line each and the turn may disappear. That's fine — that's a Short. At 1,200+ words the full shape has room to breathe. Same template, different scale.

The Shorts variant

SHORT: [working title]   Target: 150–225 words (60–90 sec)

[HOOK] 10–15 words. One sentence.
[MECHANISM] One line. Not a section.
[PAYOFF] Close the loop. Fast.
[BEAT] Does this stand alone without the long-form?

Every script on this site is turned into video with RevID.ai

Same tool for every case study documented here. 20% off with code THEAIBUILDLAB.

Try RevID.ai →
Part 9

Mistakes and quick reference

MistakeConsequence
Prompting before you have a conceptA competent summary nobody stays for.
Treating hook templates as the answer73.5% of hooks fit no formula. Templates are a fallback, not a target.
Writing a rhetorical question as a hook"Ever wonder why?" is filler, not an open loop.
A hook longer than a sentenceRevID's math: 10–15 words of a ~120-word script.
Underwriting the mechanismFacts in a pile don't retain. A causal chain does.
No turnThe script flattens out. The turn is the retention beat.
Never closing the loopIf the hook asked something, answer it.
A summary instead of a takeawayReframe, don't recap.
Cutting to fit a template length90+ seconds is the largest bucket at 35.7%. Sub-30 is 16.3%.
Planning a calendar on an estimated runtimeMeasure your own render once — words-per-minute varies by voice and settings.
Compressing a long-form script into a ShortA Short is one beat, not a summary.
Static captions, or none86.6% caption, and 96.8% of those animate. It's a retention mechanism.
Long clauses in a captioned scriptWord-by-word reveal turns them into a crawl.
Publishing the AI's first draftSame principle as the first render. The script is where originality lives.
Not diffing against your outlineThat's where added inaccuracies hide.
Never reading it aloudFree, and it catches what silent reading misses.

Quick reference

Hook — question12.2% (most common classified)
Hook — bold statement6.6%
Hook — storytelling5.4%
Hook — statistic2.2%
Hook — other / mixed73.5%
Hook length10–15 words of a ~120-word script
Average RevID script117.5 words
Average exported video115.5 seconds
0–15 sec4.4%
16–30 sec11.9%
31–60 sec31.1%
61–90 sec16.9%
90+ sec35.7% — largest single bucket
Sub-30 total16.3%
Captions enabled86.6%
Animated, of captioned96.8%
Vertical 9:1677.9%
Faceless AI voiceover55.3%
Our measured pacing880 words / 352 sec = 150 wpm exactly
Words for 60 sec150
Words for 10 min1,500
StructureHook → mechanism → turn → payoff → takeaway
Beat checkCan each mechanism step stand alone as a Short?
Sources & caveats

Primary data

  • RevID.ai, "The Anatomy of a Viral Short: What 756,851 AI-Generated Videos Tell Us About Content in 2026," by Tibo, published March 18, 2026 (revid.ai/blog/the-anatomy-of-a-viral-short-2026). Source for all hook classification, duration, caption, narration-style, and aspect-ratio data in Parts 2, 4, 5, and 6. Methodology: hook classification via heuristic NLP classifier on 348,346 scripts with usable text; duration from Amplitude tracked render events; 756,851 total export records covering 2025, 429,105 of them from 15,919 unique creators in 2025 alone.

Our own material

  • "Why You Always Think You're a Better Driver Than Everyone Else" — full script, 1,187 words, real published video and companion case study. Section markers and structure used throughout Parts 2, 3, and 8.
  • "Why Losing $20 Feels Worse Than Finding $20 Feels Good" — 1,101 words, used in Part 3 as a second structural example.
  • "Why Online Arguments Escalate Faster Than In-Person Ones" — 1,038 words, used in Part 3 as a third structural example.
  • RevID render test: an 880-word test script rendered at exactly 352 seconds — 150 words per minute, the pacing figure used throughout Part 4 and this site's credit calculator.
  • Script workflow (Part 7): research first, outline the accurate skeleton, expand via RevID, then read the generated script once checking for anything added or reworded that isn't in the outline.

Carried forward from earlier guides in this series

Caveats

The RevID data describes RevID's own creator base specifically — genuine first-party data with a disclosed methodology, but not independently audited, and the report itself states it may not generalize to all short-form video. Hook classification is heuristic, not human-reviewed, so the 73.5% "other/mixed" category likely contains both genuinely novel hooks and classifier misses — we can't distinguish those from the published data. The 150 wpm pacing figure is measured from our own renders, not RevID's platform average, and will vary by voice, style, and settings. This is not a guarantee that following this structure produces a specific outcome — RevID's own FAQ states no tool can guarantee views, and neither can a structure.