Somewhere in your reading life there’s a book you finished in two days, thought about for a month, pressed on four people, and never logged anywhere. There’s also a book sitting on your want-to-read shelf that you actually started, disliked, and quit on page sixty, where it will sit listed as an intention for the rest of time.
Both of those are the same problem pointed in opposite directions, and every recommendation system you use is built on top of them.
Four datasets, not one
What you read. What you log. What you rate. What you display publicly. Social reading platforms treat those as a single record, and they are not. They’re four datasets of decreasing size and increasing self-consciousness, and almost everything a system could learn about your taste lives in the gaps between them.
The gaps aren’t dishonesty. I want to be careful here, because the lazy version of this argument is that readers perform their taste and lie on the internet, and while a little of that is true it isn’t the interesting part. The interesting part is that the gaps are structural. They’d exist even among people trying scrupulously to be accurate.
Logging asks for a deliberate act at the worst possible moment
You finish a book at eleven at night. What you want to do next is sit with it, or start the next one, or go to sleep. What the platform wants is for you to open an app, find the title, mark it read, choose a star value, and consider writing a review.
So logging happens in batches, from memory, weeks later. That’s not a record, it’s a reconstruction, and reconstructions are shaped by what was memorable rather than what was read. The book that consumed a rainy Sunday and left nothing behind is exactly the book that doesn’t make it into the batch.
Then there are the books people leave out on purpose. The romance devoured in three days during a bad week. The self-help book nobody wants sitting on a public shelf under their real name. The fourth reread of a comfort novel, which most platforms can’t even represent, so the single most emphatic thing a reader can say about a book, which is that they went back to it again, has nowhere to go.
And abandoned books have no honest home at all. Marking something did-not-finish means making a small public verdict about a book you might simply have met in the wrong month of your life. It’s easier to leave it on want-to-read, where it quietly reads as an intention instead of what it really was, which is a genuine attempt that didn’t take. We’ve written before about where readers actually quit, and the quitting points are enormously informative. They’re also almost entirely unrecorded.
The stars aren’t doing much work either
Even when a book does get logged and rated, the rating carries less than you’d think. One independent analysis of Goodreads rating data found that around 80% of adjusted book ratings fall between 3.5 and 4.2. A scale where nearly everything lands within seven tenths of a point isn’t really a scale, it’s a formality.
Some of that compression is positivity bias, and positivity bias is mostly rational. You rate books you chose, and you chose them because you expected to like them, so of course the distribution leans high. But it means the rating is doing very little to separate the book you thought was fine from the book you’d take to a desert island, and those two books should not produce the same recommendation.
Netflix ran this experiment already
This is the part I find genuinely clarifying, because someone spent an enormous amount of money learning it.
Netflix had a five-star system and one of the most famous recommendation engines ever built. What they found was that people rated documentaries five stars and silly comedies three stars, and then went home and watched the silly comedies. The ratings were sincere. They just described a different person than the viewing history did.
A rating is a statement about the reader you’d like to be. Behavior is a statement about the reader you are on a Tuesday night.
Netflix replaced stars with thumbs in 2017, and their VP of product, Todd Yellin, put the reasoning bluntly at the time: they made ratings less important because the implicit signal of your behavior is more important. The simpler control also produced about 200% more ratings, which tells you something on its own about how much friction the five-star ask was adding.
Books are harder than television in every practical respect, since there’s no single player logging your progress and the reading happens across paper, apps, and audio. But the underlying lesson transfers exactly, and most reader tools have not absorbed it.
The missing data isn’t missing at random
There’s a precise name for this problem in the recommender-systems literature, and it’s worth knowing because it explains why the gaps are worse than they look.
Almost all collaborative filtering assumes ratings are missing at random, meaning the books you didn’t rate are roughly a random sample of the books you haven’t gotten to. They aren’t, and researchers have spent fifteen years establishing how badly they aren’t. People rate what they’ve been exposed to and what they had positive feelings about, so negative and ambivalent experiences go missing at much higher rates than positive ones.
That’s the crux. A system trained only on what you logged isn’t merely working with less data, it’s working with data bent in a consistent direction, and the bend is invisible from inside the dataset. It learns the reader who shows up in public, then recommends to that reader, and the recommendations come back slightly off in a way that’s hard to articulate. They’re not wrong about the person on the shelf. They’re wrong about the person holding the book.
What the useful signals actually are
Here’s what I’d want to know about a reader, none of which is a rating, and all of which is something they already did.
How fast they finished it, because time-to-finish is the most honest rating anyone has ever produced and nobody has to be asked for it. Whether they went back to it. What they picked up immediately afterward, since the book that follows a great book is a real judgment about the great book. Where they stopped, if they stopped. What they scrolled past over and over without ever opening, because a repeated skip is a genuine opinion. Whether they tore through four books by one author in a month and then never touched that author again, which is a complete arc and says more than any single star value could.
None of that requires a form. It’s the residue of reading, and on most platforms it’s thrown away in favor of a number the reader had to stop and produce on purpose.
What we’re building toward
This gap is most of why Siftivo exists. We’d rather infer than interrogate, which is why the books you skip in your daily picks count as signal rather than noise, why we don’t gate recommendations behind rating a pile of books first, and why we care more about whether you finished something than whether you gave it four stars.
It’s the same instinct behind the Reading Index, which measures cultural circulation instead of units sold. Circulation means being read and argued about and handed to somebody, and it’s a harder thing to measure than a purchase or a star, which is exactly why it’s worth measuring.
Your public shelf is a decent portrait of the reader you’d like to be seen as, and there’s nothing wrong with wanting that. It’s just a poor portrait of the reader you are at eleven on a Tuesday night deciding whether to start something new.
That second reader is the one who needs a good recommendation, and almost nobody is building for them. Come see what we’re doing.
