Between early 2023 and early 2026, Amazon's self-published catalog grew 38.3 times over. Revenue during that same window rose just 8.9 times. That gap is the entire story for anyone earning a living with words: vastly more books are competing for a pot of money that, by comparison, barely moved.
A fresh analysis covering 14,419 randomly sampled self-published e-books argues that AI-generated titles are not succeeding because they're better. They're succeeding because there are so many of them. And the books absorbing the damage are the ones where no AI text was detected whatsoever.
A different approach to measuring the problem
Earlier efforts to size up the problem tended to sample brief book previews and estimate from there. This study classified each book on its complete text using the Pangram v3.3 detector, which its developers say has a false-positive rate of 0.04 percent. Titles were sorted into three groups according to how much text was flagged as machine-written: none, light at up to 25 percent, and substantial above 25 percent.
The sales side drew on an internal dataset maintained by one of the five major US publishers. It follows roughly 500,000 Amazon titles and, per the researchers, accounts for around 95 percent of all e-books sold on the platform each day.
Skim the topline numbers and AI books appear to be failing. Titles carrying substantial AI content account for 20 percent of the studied catalog yet capture only 12.1 percent of sales and 11.3 percent of revenue. Books with no detected AI text represent 62.9 percent of the catalog and collect 72.5 percent of revenue.
Human-written books are losing ground on their own
The count of titles selling each quarter climbed 19.2 times, set against that 8.9x growth in revenue. Every slice shrank - including those belonging to writers doing the work themselves.
Measuring 2023 releases against 2025 releases across an identical post-release window, revenue per book declined in six of eight genres. Restrict the view to books with no detected AI text and the picture darkens: revenue slipped in seven of eight.
That result dismantles the convenient counterargument that a wave of worthless AI titles is dragging the average down. Human-written books are pulling in less money on their own terms. The authors label the phenomenon "dilution," while emphasizing that their comparisons are observational and associational rather than experimental proof of causation.
One genre went the other way. Fantasy/Supernatural/Horror, where AI text arrived last and gained the least traction, saw revenue per book rise 35 percent for titles with no detected AI text. The researchers cite that exception as evidence against attributing the trend to some general market downturn.
Bestseller rankings and the accounts behind the flood
Fresh Top 25 entries containing substantial AI content rose from close to zero to 31 percent over the study window. Churn at the summit accelerated. The proportion of books with no detected AI text keeping a Top 25 slot from one quarter into the next dropped to roughly 28 percent at one stage before finishing near 62 percent.
Thousands of hobbyists aren't behind this. Among 385 author identities that released more titles with substantial AI content following their first AI book, 287 increased their monthly output afterward. Gross revenue before platform fees for the highest-earning pseudonym reached $1.7 million across eight titles. The single top-grossing AI book earned $643,000 from 80,431 copies sold.
The rare-phrase test and what it means for the courts
To measure how far the language in successful AI books overlaps with existing works, the researchers pointed the Allen Institute for AI's infini-gram tool at the Google Books index. They searched for rare expressions appearing in five or fewer Google Books volumes and completely missing from a 4.7-trillion-token web snapshot. Wording that distinctive implicates published books, not the open internet.
Across the 50 top-grossing titles with substantial AI content, such rare expressions accounted for 45 percent of the text. The figure was 37.7 percent for the top 50 books with no detected AI text, and 19.1 percent for award-winning or award-nominated fiction. Inside the AI books, overlap rose 7.6 percentage points with every tenfold increase in revenue.
Speaking to Ai2, researcher Tuhin Chakrabarty cautioned that an AI detector delivers only "an estimate - a score for how likely a passage is to be synthetic," with no indication of where the language originated. But when a questionable text also contains rare expressions missing from the web and present in only a few books, one "can say with some confidence that it was taken from books."
That kind of evidence, he said, "acts as circumstantial evidence that supports an AI detector score" and also "helps debunk some hackneyed arguments that liken human reading of books to AI being trained on books" - a defense AI companies routinely reach for in copyright disputes.
Judge Vince Chhabria sided with Meta in Kadrey v. Meta in June 2025, but paired the ruling with a sharp caveat. He said he found it hard to imagine that feeding copyrighted books into a product that generates billions in revenue, while spawning a potentially limitless flood of competing works, would count as fair use. The plaintiffs in that case brought no empirical evidence of market dilution. This study supplies exactly what was missing: books with no detected AI text making less money as AI titles flood in.
Amazon keeps the data private
Anyone publishing through Kindle Direct Publishing is required to disclose AI involvement. Amazon does not relay that disclosure to shoppers. The one party holding a clean record of which books are machine-written keeps it in-house while customers browse blind.
For years the platform has struggled with AI titles appropriating the names and styles of established authors, and its principal remedy has been a limit of three publications per day. The language-overlap results align with a November 2025 study showing that language models can reproduce passages from copyrighted books almost verbatim, and with a second fall 2025 paper finding that two books are enough to fine-tune a model on an author's style.
Why this matters for writers
The practical implications for writers are straightforward. This study strips away whatever comfort the "AI books aren't selling" narrative provided - plain-language books sold less per title in seven of eight genres even before accounting for AI competition.
The numbers also offer a concrete piece of evidence if you're considering legal action. Judge Chhabria came close to spelling out exactly what a plaintiff needs, and this study provides the market-dilution evidence his ruling said was missing. If your per-title earnings have sunk since 2023 without you changing output or quality, the data now backs up the suspicion. As Chakrabarti put it in a follow-up paper, "this is no longer a matter of someone's anecdote about how they're earning less" - dilution has all structural drivers.
Your membership also unlocks: