Book discoverability is not three separate battles fought with three different tools. It is one metadata layer, read three different ways, and won or lost before promotion ever starts. This post is the entry point to four decisions covered later: the keyword boxes, fiction comps, non-fiction comps, and answer-engine readiness. It follows Before You Publish: The Decisions That Are Hard to Undo, which covered the pre-publication inventory, including BISAC and Thema classification codes and subtitle assessment. This piece picks up where that inventory ends and shows what each discovery surface actually rewards.
Book discoverability works because the same fields that feed a retailer’s search algorithm also feed a search engine’s query matching and an AI model’s summary of what the book is. When those fields agree with each other across every platform, all three systems reinforce the same signal. When they contradict each other, each surface sees a weaker, less confident version of the book. The sections below show what each surface rewards, what the evidence actually supports, and how two worked examples put the same method to use on fiction and non-fiction.
Three Surfaces Read the Same Pile of Text for Book Discoverability
Book discoverability is won or lost in roughly two hundred words of author-supplied text plus a handful of classification codes. That pile is small and specific: title, subtitle, description, categories, the seven KDP keyword boxes, and comp titles. Categories get full treatment in a later pillar and appear here only as part of the wider pile.
Splitting this work across specialists tends to backfire. A keyword consultant, an SEO writer, and an AI-visibility consultant each rewrite one fragment in isolation, and the assertions about the book stop agreeing with each other. Each discovery surface reads one short document the author writes once, just through a different lens.
According to Ingram Content Group, subject categories and keywords are the two most important discoverability elements a publisher controls. Timing compounds this further: the IBPA’s trade data shows metadata delivered well ahead of publication correlates with stronger sales, which places this work before launch rather than after it. The inventory work covered in Before You Publish, the BISAC and Thema codes and subtitle assessment, is the raw material this pillar puts to work.
What Each Surface Actually Rewards
The same text gets scored three different ways, so the useful question is what each scoring system is actually looking for. Amazon is scored on behavior, Google on relevance to a query, and answer engines on whether a book can be described consistently by more than one source.
Amazon rewards conversion signals on fixed slots
Amazon’s search behavior is inferred by practitioners and documented almost nowhere by Amazon itself.
- Fixed slots: Title, subtitle, description, seven keyword boxes, and categories are the only fields an author controls directly.
- Behavior on top: Click-through and purchase behavior appear to reweight those slots over time.
- Test it: Search your own keyword strings in an incognito window and record which titles surface.
Google rewards the better answer to a query
A book’s Amazon page competes with review posts, retailer pages, and aggregators for the same search query.
- Query match: The page answering the searcher’s actual question ranks, not the page that repeats the title.
- Owned pages: An author-controlled book page can outrank a retailer listing for long-tail queries.
- Test it: Search the three phrases a reader would plausibly type and see who is already answering them.
Answer engines reward a describable entity
Large language models summarize a book from whatever assertions about it they encountered during training or retrieval.
- Consistency: The same genre, premise, and comparison language should appear across retailer, author site, and catalogue records.
- Attribution: A book that only exists on one listing has one unverified source behind it.
- Test it: Ask three different assistants “what is [title] about” and compare the answers to your description.
What the Evidence Supports and What It Does Not
Amazon documents almost nothing about how categories, keyword boxes, or search ranking actually work. Most advice circulating in author communities is practitioner inference dressed up as certainty. It’s worth holding that loosely, even when the advice sounds confident.
The keyword evidence deserves careful handling. The eBOUND Canada / Ontario Creates study found a significant increase in discovery, measured as views on Amazon, with human-created keywords producing the largest increase. The same study found that keywords “had no effect on sales regardless of the ebook’s category, sales ranking or publication year.” Keywords are demonstrated to increase views. They are not demonstrated to increase sales.
A pattern that shows up often: an author packs the keyword boxes with genre words, watches views climb, and assumes sales will follow on their own. When they don’t, it’s tempting to blame the algorithm rather than look at the description and comps sitting beneath those keywords. That gap, between being found and being chosen, is exactly what the study above measured.
Two limits apply every time this study gets cited. It tested ONIX keyword fields supplied by publishers through distribution, a different mechanism from KDP’s seven keyword boxes filled in directly by an author. It also covered 600 Canadian ebooks over a single quarter. The study’s own authors say converting views into purchases needs further research, not that it doesn’t happen.
For proportion, word-of-mouth led discovery at 36 percent in the 2025 Canadian Leisure and Reading Study, ahead of bookstores, libraries, social media, and online retailers. Personal recommendation led similar surveys more than a decade earlier, according to Pew Research Center. Metadata doesn’t replace human recommendation. It just decides what a curious reader finds once someone else has already pointed them your way. Build a testing habit instead of chasing certainty: change one field, record the date, and watch views rather than rank.
Key Evidence:
- Titles with complete basic bibliographic records sell substantially more than titles missing that information, per the Independent Book Publishers Association.
- Human-created keywords produced the largest gains in Amazon views but showed no measurable effect on sales, per the eBOUND Canada / Ontario Creates study of 600 Canadian ebooks over one quarter.
- Word-of-mouth led book discovery at 36 percent, ahead of bookstores, libraries, social media, and online retailers, per the 2025 Canadian Leisure and Reading Study.
- Readers without a known author to search rely on topical browsing and similarity-based tactics, per the Journal of the Association for Information Science and Technology.
Worked Examples: Two Books, One Pile of Text
Two full sample reports, published at Perivane, show the same method applied to fiction and non-fiction. Agatha Christie’s The Murder at the Vicarage is a cozy village mystery with an amateur detective. Its keyword phrasing and comp titles describe reading experience, setting, and detective type, the same phrases that serve a Google long-tail query and give an answer engine consistent assertions to repeat.
J.C. Ryle’s Expository Thoughts on Matthew needs a different vocabulary entirely. As a devotional commentary, the searcher’s problem is often stated directly in the query itself, and its comps signal reading level and theological tradition rather than mood. Fiction comps answer “what will this feel like.” Non-fiction comps answer “will this solve my problem, and at what level.” You might notice this yourself if you compare the two: fiction and non-fiction fill the same seven boxes with completely different kinds of language, and every surface reads both.
This positioning work matters most for authors without name recognition. According to research published in the Journal of the Association for Information Science and Technology, readers without a known author to search fall back on topical browsing and similarity-based tactics. Comps and categories substitute for the author-name search these readers can’t perform.
The Four Decisions This Pillar Covers Next
Think of this as a map, with each item stated as a decision rather than a tactic. Book discoverability improves when these four are made together, not in isolation.
- Decision one, the seven keyword boxes: What goes in them, why phrase-level entries behave differently from single words, and how to run a controlled test on your own listing.
- Decision two, fiction comps: How to choose comparison titles that describe reader experience honestly, and why over-reaching comps damage both conversion and the assertions an answer engine inherits.
- Decision three, non-fiction comps: How to pick comps that signal problem, level, and tradition, using the Ryle example as the pattern. See how readers discover self-published novels for the fiction side of this comparison.
- Decision four, answer-engine readiness: What it takes for a book to be a describable entity, with consistent claims across retailer listing, author page, and catalogue records.
Categories, meaning BISAC, Thema, and retailer placement, belong to the next pillar and aren’t settled here. Four decisions, made once, feed three surfaces. No amount of promotion repairs them afterward. Doing all four properly means reading comp descriptions, checking live search results, drafting and redrafting phrasing, and keeping one master record so the fields stay consistent across formats and retailers.
Why Book Discoverability Matters Now
Readers increasingly arrive through summarizers and recommendation surfaces that never show the author’s page directly. An author with no existing audience depends entirely on metadata to make the introduction. Metadata written once keeps working for the life of the backlist, in the right direction or the wrong one, which makes getting it right early worth the hours it takes.
Conclusion
Book discoverability is one problem with three readers, and the pile of text involved is small enough that one person can get it right. Metadata reliably affects whether a book is seen; the link from being seen to being bought is weaker than commonly claimed, and it’s worth remembering that before promising results to yourself or anyone else. There’s no single right way to sequence this work, only the four decisions above, made deliberately instead of by accident.
None of this is quick. Comp research, keyword drafting, live search checks, and cross-platform consistency take hours per title, repeated for every book and every format an author publishes. The two worked examples above, The Murder at the Vicarage and Ryle’s Expository Thoughts on Matthew, are published in full at perivane.com for anyone who wants to see the finished output of that work before attempting it themselves.
Start with the keyword boxes post next, then move to whichever comps decision matches your book’s type, fiction or non-fiction. The four decisions ahead build directly on the metadata pile covered here.