Skip to content

Methodology

How we counted, and where the counting breaks

The articles on this site make numerical claims about a whole genre. This page is the part that lets you decide whether to believe them.

Last updated 4 August 2026

The corpus

One dataset sits behind every figure on this site: 65,510 books, each appearing exactly once, assembled from Goodreads in 2024 by two routes.

The bulk of it, 53,561 books, came from looking up individual titles seeded from a romance catalogue, one book at a time. The remaining 11,949 came off 10 large Goodreads romance lists: “Best Book Boyfriends”, “Best Ever Contemporary Romance Books”, “Favorite Historical Romance Novels” and similar.

That split is worth knowing, because the two routes are biased differently. A themed list is a popularity artefact and a taste artefact at once, since it holds the books enough readers voted for. A title lookup carries whatever bias the seed catalogue had, but nobody voted. The list-sourced portion is also where most of the non-romance crossovers arrive, because readers put all sorts of things on a list called “Best Love Stories”.

It is a snapshot rather than a live index. Nothing published after 2024 is in it, it has not been topped up since, and the articles say “2024” wherever the recency matters.

For each book the corpus holds the bibliographic basics (title, author, series, publication date, average rating, rating and review counts), up to 30 reviews, and fourteen prose descriptions covering different facets of the book: genres, themes, plot, setting, character archetypes, main character, writing style, emotional impact, strengths, genre-specific craft, steaminess, conflict, weaknesses and content warnings.

Where the facet descriptions come from

Those fourteen paragraphs were not written by hand and were not scraped from anywhere. They came out of an automated reading pass over each book's publisher description and its reviews: software read the available text about a book and wrote a paragraph per facet. We are being plain about that because it sets a hard ceiling on what any of these numbers can mean.

These are descriptions of descriptions. The steaminess paragraph is not a count of scenes in the novel. It is an account of how the book presents itself and how its readers talk about it. That happens to be the signal you work from when you browse a shelf, which is why it is useful at all, but it is not the same as having read the book.

What the corpus is skewed towards

It is heavily recent. 48,872 of the 65,510 books were published in 2020 or later, which is roughly 75% of the whole thing.

Corpus composition

Three quarters of the corpus was published since 2020

Books by five-year publication window, 2000 onwards.

2000–04
819
2005–09
1,750
2010–14
6,127
2015–19
5,393
2020–24
48,872
Older books are present only if they still appear on a Goodreads romance list today, which selects for the ones that lasted. Any claim about change over time has to carry that caveat, and ours does.
Show the numbers as a table
PublishedBooks
2000–04819
2005–091,750
2010–146,127
2015–195,393
2020–2448,872

It is also skewed towards books that got attention. Self-published titles with a handful of ratings are under-represented relative to how many of them exist, which is a real gap in a genre where the indie long tail is most of the publishing.

And the edges are fuzzy. 59,108 of the 65,510 books describe themselves in genre terms that include romance. The rest arrived because a reader put them on a romance list. Some of those are crossovers with a strong romantic thread and some are simply books somebody loved. We kept the whole set rather than drawing a line we would then have to defend, and we say “romance and romance-adjacent” where the distinction matters.

Turning prose into counts

The facets are paragraphs, and paragraphs cannot be added up. Three separate methods turn them into something countable, and each fails differently.

1. Quoted labels

The character-archetype paragraphs put archetype names in quotation marks, the “wounded healer” and the “chosen one” and so on. Pulling out every short quoted phrase gives a vocabulary of 18,129 distinct labels across 55,165 books (84% of the corpus).

This method is clean, but it measures the describer's vocabulary as much as it measures the books. The most common label in the entire corpus is wounded healer, and terms like trickster and seeker rank far higher than any romance reader would put them. That is a Jungian habit in the describing software rather than a fact about romance novels. We use this vocabulary to show what language gets used, never as a count of what is in the books.

2. Term matching

Tropes, subgenres and character types are counted by matching a fixed, hand-written list of phrasings against each book's facet text. “Enemies to lovers” and “enemies-to-lovers” count together, “fake dating” sits alongside “pretend relationship”, and so on. Every pattern is deliberate and none of them are generated.

The obvious flaw: a match proves the phrase appears, not that the book has the thing.“Refreshingly free of the usual love triangle” contains “love triangle”. Rather than tell you that this is rare, we went and measured how often it happens. Those numbers are further down.

3. The heat rubric

Heat level comes from the steaminess paragraph, which states a level in a near-formulaic opening sentence (“contains a moderate level of romantic and sexual content”). The classifier reads that first sentence, collects every heat word, discards the ones sitting behind a denial, and takes the highest surviving level.

That denial step is doing almost all of the work. Our first version didn't have it, and the result was wrong in a way that was hard to miss once we looked: it rated inspirational Christian romance 38% explicit. The reason is that “the author avoids explicit descriptions” contains the word “explicit”, so the classifier was reading denials as confirmations. With the fix in, inspirational romance comes out at 0.4% explicit, which is about what anyone who reads them would tell you.

We are publishing that mistake because it is the kind that never announces itself. Every number in the broken version looked perfectly reasonable on its own. It only fell over on a subgenre where we already knew the answer, which is an uncomfortable thing to notice about your own pipeline.

The classifier places 64,137 of the 65,510 books. The other 2.1% describe intimacy without ever stating a level, and they are excluded from the heat figures rather than guessed at.

The published charts use four bands. The underlying rubric distinguishes five, but the corpus's own wording doesn't draw a reliable line between “steamy” and “explicit”, so those two get merged instead of being reported as a distinction we can't support.

What we measured about our own accuracy

Hand-checking the heat classifier

Three random samples of books were pulled and read by hand against the classifier's output. The first two were used to find and fix faults: the denial problem above, plus a tendency to get swayed by heat words further down the paragraph that describe one scene rather than the book.

The third sample of 40 books was drawn after the fixes and used only to measure. On that sample, 39 of 40 books came out in a band a human reader would defend, and the disagreements that are left sit at the steamy/explicit boundary. Which is why those two bands are reported merged.

How often a term match is really a denial

For the term-matching method we sampled matches across the corpus and checked how many of them sit behind a negation: a “no”, an “avoids”, a “subverts”, a “refreshingly free of”. Overall 2.9% of sampled matches are negated, so the trope counts on this site are overstated by roughly that much.

Error measurement

Share of matches that are actually a denial of the trope

Sampled matches per trope, checked for a negation immediately before the phrase. Lower is a cleaner count.

insta-love
15.4%
love triangle
8.0%
fake relationship
2.9%
enemies to lovers
1.8%
slow burn
1.2%
grumpy / sunshine
0.9%
forced proximity
0.8%
second chance
0.5%
Sampled at roughly one book in seven, up to 400 matches per trope. This is a floor, not a ceiling: it catches negations sitting close before the phrase and will miss ones expressed further away or across a sentence boundary.
Show the numbers as a table
TropeMatches sampledNegatedNegated %
insta-love1562415.4%
love triangle225188%
fake relationship24472.9%
enemies to lovers40071.8%
slow burn40051.2%
grumpy / sunshine10810.9%
forced proximity36630.8%
second chance40020.5%

This matters most for the comparisons the articles actually make. Because the error is roughly similar across tropes, ratios between tropes survive it better than absolute shares do. Still, read any share printed here as “about this much, slightly over” rather than as a precise figure.

The association measure

Where the articles say two things turn up together “more often than chance”, the measure is books containing both, divided by the number you would expect if each were distributed independently across the corpus. One times is exactly chance. Pairs are only reported above a minimum co-occurrence count, 150 books for tropes and 200 for character types, so that a ratio can't be manufactured out of a handful of books.

This measures association. It does not measure causation and it says nothing about authorial intent. It also inherits every flaw above: if two tropes are both over-counted by denials, their association is measured on inflated numbers.

Claims this data cannot support

  • Anything about the whole genre. This is one popularity- weighted snapshot of Goodreads, not a census of romance publishing. The long tail of self-published romance is thinly represented.
  • Clean historical trends. Older books survive into this corpus only by still being listed today. A trend line here mixes real change with survivorship.
  • Anything about quality. The corpus carries ratings, and we deliberately do not build claims on them. Rating differences between tropes are small, confounded by subgenre and audience, and not worth the sentence they would take.
  • Same-sex romance at any depth. The lead-pairing analysis relies on gendered role language in the source text and is largely blind to queer romance. That is a limit of the extraction method, and we say so where it appears.
  • Anything about a specific book. These are aggregate counts. A single book's facet paragraph can be wrong, and at this scale that washes out of a percentage while remaining very much present in any individual row.

Reproducing it

The analysis is three scripts: one streams the corpus and extracts structure, one computes every figure and writes the data files this site renders from, and one runs the error measurements above. The site reads those files directly, so no number here is typed in by hand. If an article disagrees with the data, the article is being generated wrong rather than the data being stale.

If a figure here looks wrong to you, it may well be. Write to hello@smittenspines.com and say which one. Corrections get published on this page.

Corpus summary: 65,510 books · 64,137 heat-classified · 39 tropes tracked · 25 character types tracked · 21 subgenres. Back to the reading room.