The PreList Industry Intelligence
Methodology.
Every published figure should survive a sceptical reader. This page says where the data comes from, how it is measured, what is excluded and why, and what the numbers cannot claim.
The corpus.
The research corpus is a library of professionally produced, publicly released screenplays: films and television episodes whose scripts studios and publishers have made available. Each new release requires recorded source provenance and a metadata review covering year, format, writers and primary genre. Reviews are tied to the held file bytes and the exact measurement result; changing either requires a new review. Writing credits describe screenplay and revision writers on the held draft, excluding literary and story-only authors. When a cover omits an author, the review explicitly identifies the external credit source used. Earlier releases predate that review ledger; missing historical provenance and incomplete credits are disclosed in their packages. Each script has one primary genre from a fixed list of ten.
We publish statistics derived from these scripts. We do not republish the scripts themselves, and no dataset we release contains screenplay text.
How screenplays are measured.
Screenplay structural statistics are counted deterministically from the screenplay text by our measurement engine: page counts, scene headings, scene lengths in eighths of a page, locations, speaking characters, dialogue and action volume. Run the engine twice on the same script and it returns the same numbers. PDF page counts include title pages, blank pages and other front matter in the held file; they are not counts of numbered story pages. Files with unseparated prose or artwork appendices are excluded from the reviewed benchmark. Repeatability does not establish that the extraction is correct. No LLM judgement scores are included. Runtime is a fitted estimate, explained separately below, not an observed film running time.
The engine is versioned (current studies use v1.5), and each measurement records how confidently the script’s layout was read. The cohort excludes low-confidence parses. Family-specific guards remove flagged measurements from affected statistics; accepted measurements can still contain extraction errors.
Runtime is a modelled estimate.
The archived runtime field estimates minutes from script composition. For scripts of at least 60 pages, engine 1.5 uses:
minutes = 41.6 + 3.701 × (dialogue words / 1,000) + 2.75 × (rendered action lines / 100)
The result is rounded to a whole minute. Scripts below 60 pages use one minute per measured page. The fitted coefficients came from 21 films in the earlier calibration corpus; the same coefficients remain in this engine version. They assume a stable relationship between extracted script composition and released film length.
That small historical calibration is not an independent validation on this published cohort. Draft changes, performance, editing, extraction errors and changing measurement definitions can all affect the estimate. We have not established its prediction error for these 428 scripts. Runtime is retained for transparency in the frozen package, but is not a headline benchmark or a claim about the running time of any particular film.
What the figures mean.
- A median divides the eligible scripts in half. The middle 50% runs from the lower to the upper quartile. It is not a confidence interval, a quality score or a target for a draft.
- Each distribution counts scripts. Scene-length distributions summarise each script’s median scene length, rather than pooling every scene. Shares of interior scenes, daylight scenes or dialogue are also calculated within each script before being summarised.
- Dialogue:action compares rendered lines, not words or screen time. A speaking character is a detected character cue; a monologue is a dialogue block longer than six rendered lines. Neither definition establishes narrative importance.
- Locations are normalised base names extracted from headings, not shooting sites. Similar names can be grouped, and naming variants can still split one place.
Cohorts, floors and exclusions.
Each study defines its cohort before any number is computed. The current flagship cohort is: Produced feature screenplays released 2000-2026, measured from the file we hold, excluding parses the engine rated low confidence.
Features and television scripts are never pooled into one figure. No genre cut is published below n=20: eligible samples under the floor are counted and shown, but their statistics are held back. This is a publication rule, not proof that a sample above the floor represents the wider industry.
Of 557 scripts in the corpus, 428 are in the current flagship cohort and 129 are excluded, each with a stated reason. The most common:
- 49 · format is TV Pilot, not Feature
- 38 · year outside 2000-2026
- 19 · format is TV Episode, not Feature
- 6 · parse confidence low
- 5 · not measured with engine 1.5
- 2 · no screenplay file
The complete exclusion list, with every title named, ships inside each dataset’s manifest.
What this research never uses.
Screenplay studies use professionally produced, publicly released scripts. Market studies use the public sources described on their own pages. Scripts uploaded to The PreList by writers are never part of public research, in any form, aggregated or otherwise.
This is the same boundary the product itself keeps: your work is yours, and it is not training data and not research material.
Reproducibility.
Screenplay studies use frozen data; the cost-of-entry study explains its live directory calculation separately. Each screenplay study renders from a frozen snapshot that records its cohort rules, every inclusion and exclusion, the engine version and a checksum of the data. When the corpus grows or the engine changes, a study updates by moving to a new dated snapshot, and its update history says so. Publication requires an explicit approved version and checksums are verified on load and build.
Browse the dated release history and public manifests. The current flagship snapshot is anatomy-of-a-produced-screenplay-2026-09-v5, generated 20 September 2026.
When a measurement fails.
Screenplay PDFs are not uniform, and extraction can fail for one part of a script while succeeding for the rest. A script whose scene headings are set in a way our parser does not recognise may still yield usable dialogue counts; a scene flag alone does not establish whether another measurement is accurate.
We apply plausibility thresholds to each family of measurements. A feature averaging more than five pages per scene, or containing fewer than twenty speech blocks, may indicate missed headings or dialogue. These are heuristics for likely extraction problems, not mathematical impossibilities. A failed family is excluded from measures that depend on it; other eligible measures remain. Each chart reports its own resulting sample size.
The 20 September source review withdrew the earlier dialogue exemption for No One Will Save You: its apparent speakers were action directions. Shared cue checks now flag prose-like leading cues, unresolved name variants and substantial unattributed speech. Source-reviewed failures, including lyrics mistaken for speakers, receive additional family exclusions. These conservative checks can also withhold valid material; they do not estimate an error rate. The package includes the thresholds, flags and review reasons. Of the 428 scripts in the current cohort, 200 carry at least one flag, and every one of them is named with its reason in the snapshot manifest, alongside the dataset.
Limitations.
- Publicly available produced scripts skew toward acclaimed and awards-circulated titles. This cohort describes produced screenwriting as it survives publication, not a random sample of everything shot.
- Available drafts vary: some are shooting scripts, some earlier drafts. Where a film differs from its published script, we measured the script.
- Some character cues can be misread as prose, and name variants can split one character into several. Current guards do not detect all such cases. The current source review checked all held-file identities, reran available measurements and compared selected source pages, including all previously flagged titles. It does not provide a whole-script gold standard or a measured extraction-accuracy rate. Counts remain automatic measurements of a selected corpus.
- Screenplay measures describe association, not cause. Nothing here claims that matching a median makes a script better, and no figure on this site is a rule.
Corrections.
Corrected studies use a new snapshot with a dated note. Previous snapshots are retained unchanged. If you believe a figure is wrong, write to enquiries@teamprelist.com and include the snapshot id from the page footer.
Citation and licence
Figures, charts and datasets are licensed CC BY 4.0: reuse them, including commercially, with attribution to The PreList and a link to the study you drew from. Each study ends with a ready-made citation and a dataset request form; the download link arrives by email.
