The PreList

The PreList Industry Intelligence

Methodology.

Every published figure should survive a sceptical reader. This page says where the data comes from, how it is measured, what is excluded and why, and what the numbers cannot claim.

The corpus.

The research corpus is a library of professionally produced, publicly released screenplays: films and television episodes whose scripts studios and publishers have made available. Each script’s file source is recorded at the moment it enters the corpus, along with its year, format and a single primary genre from a fixed list of ten, so that a “horror median” means the same thing in every study.

We publish statistics derived from these scripts. We do not republish the scripts themselves, and no dataset we release contains screenplay text.

How screenplays are measured.

Every statistic on this site is counted deterministically from the screenplay text by our measurement engine: page counts, scene headings, scene lengths in eighths of a page, locations, speaking characters, dialogue and action volume. Run the engine twice on the same script and it returns the same numbers. Nothing here is generated, estimated by a model, or scored by opinion.

The engine is versioned (current studies use v1.5), and each measurement records how confidently the script’s layout was read. Scripts whose files cannot be read reliably are excluded rather than approximated.

Cohorts, floors and exclusions.

Each study defines its cohort before any number is computed. The current flagship cohort is: Produced feature screenplays released 2000-2026, measured from the file we hold, excluding parses the engine rated low confidence.

Features and television scripts are never pooled into one figure. No genre cut is published below n=20: cells under the floor are counted and shown, but their statistics are held back until the sample can carry them.

Of 357 scripts in the corpus, 268 are in the current flagship cohort and 89 are excluded, each with a stated reason. The most common:

  • 28 · format is TV Pilot, not Feature
  • 15 · format is TV Episode, not Feature
  • 4 · parse confidence low
  • 4 · year 1994 outside 2000-2026
  • 3 · year 1999 outside 2000-2026
  • 3 · year 1995 outside 2000-2026

The complete exclusion list, with every title named, ships inside each dataset’s manifest.

What this research never uses.

Our research is built exclusively on professionally produced, publicly released screenplays. Scripts uploaded to The PreList by writers are never part of public research, in any form, aggregated or otherwise.

This is the same boundary the product itself keeps: your work is yours, and it is not training data and not research material.

Reproducibility.

Published pages never read a live database. Each study renders from a frozen snapshot that records its cohort rules, every inclusion and exclusion, the engine version and a checksum of the data. When the corpus grows or the engine changes, a study updates by moving to a new dated snapshot, and its update history says so. A number on this site cannot drift quietly.

The current flagship snapshot is anatomy-of-a-produced-screenplay-2026-08-v1, generated 12 August 2026.

Limitations.

  • Publicly available produced scripts skew toward acclaimed and awards-circulated titles. This cohort describes produced screenwriting as it survives publication, not a random sample of everything shot.
  • Available drafts vary: some are shooting scripts, some earlier drafts. Where a film differs from its published script, we measured the script.
  • Screenplay measures describe association, not cause. Nothing here claims that matching a median makes a script better, and no figure on this site is a rule.

Corrections.

Errors are corrected in place, with a dated note in the affected study’s update history. If you believe a figure is wrong, write to enquiries@teamprelist.com and include the snapshot id from the page footer.

Citation and licence

Figures, charts and datasets are licensed CC BY 4.0: reuse them, including commercially, with attribution to The PreList and a link to the study you drew from. Each study ends with a ready-made citation and a dataset request form; the download link arrives by email.

Back to the research hub