A Mathematical Theory of the Antilibrary

I have too many books. This became very clear last year, when I moved.

During a long train ride, I also attempted to catalogue all the technical books I had read. I wanted to classify each one as “useful,” “evergreen,” or “garbage,” and keep track of how much I had actually benefited from it. Then, suddenly, the words of one of my favourite authors, Nassim Nicholas Taleb, came to mind: the antilibrary.

Taleb’s idea is that unread books are not useless possessions; they are a valuable reminder of everything we do not yet know. The story goes something like this, though I am unsure whether it is entirely true:

Umberto Eco famously kept a large personal library, and visitors tended to fall into two groups. Some admired how many books he had read. The more interesting visitors understood that the unread books mattered too: they mapped the edges of what he did not know.

I have always been fascinated by this concept, but none of my friends who gazed at my library seemed interested in listening to me explain it, let alone trying to formalize it mathematically.. multiple times. So here I am. Hopefully, someone will find it interesting.

(I swear I am fun at parties. Sometimes.)

1) The binary model

Suppose a library contains books. Give each book a binary reading state

Then the usual split is simple:

The realized library is and the antilibrary is . Their shares are

However, we all know that reading is not binary. I may know a reference book well without ever reading it from cover to cover. I may have technically finished a novel and remember nothing beyond the colour of its cover. The binary model throws all of that into two buckets and calls the unread bucket profound.

2) The h-index temptation

What if we take inspiration from the h-index?

The usual h-index is the largest for which a researcher has at least papers with at least citations each. It combines two quantities, the number of papers and the number of citations, by forcing them to meet on the diagonal.

We can do the same thing with books. Give each book a comprehension score and define the survival curve

is the fraction of books understood to at least level . The Library Index is

If are the scores in descending order, the same quantity is

So means that at least 72% of the books are understood to at least the 72% level. This is almost a perfect analogue of the h-index.

It is also a brutal compression. A single number discards nearly everything interesting about the shape of the collection. The two-dimensional version makes that loss visible.

Can we use two variables?

A pair naturally asks a yes-or-no question: are at least a fraction of the books understood to at least level ?

We could encode the answer as the binary function

Equivalently,

For example, would mean that at least 70% of the books are understood to at least 80%. If , fewer than 90% reach that threshold.

This is mathematically clean, but still not very expressive: almost all the interesting information sits on the boundary between and .

The feasible region

My preferred object is therefore not the binary output but the set of claims for which the output is true:

Every point represents the statement

At least of the library is understood to at least .

For the illustrative collection below:

Fraction Threshold Feasible?

The feasible set of library statements. The shaded area lies to the left of the staircase boundary.

The shaded region contains every true claim. The staircase is its frontier: points to its left are feasible, while points to its right ask the library to guarantee more books than it can at that comprehension threshold. This looks a lot like a Pareto frontier. The interior is valid, but the boundary tells us the strongest claims the collection can support.

Reading the same boundary backwards

The inverse view starts with a desired fraction of the library and asks for the highest comprehension threshold it can guarantee:

Then means that half of the library is understood to at least 82%. It is the same boundary read in the other direction: accepts a threshold and returns a fraction, while accepts a fraction and returns a threshold.

The most fundamental object is really the relation

The library decides which propositions are true. From that one relation,

and the Library Index is the furthest feasible point along the diagonal :

Calling it a fixed point is tempting, but slightly too strong for a staircase: need not equal exactly. What matters is that is the last true claim on the diagonal.

3) The utility score

We can try a totally different approach and take a page from the economists. Let’s talk about utility.

First, an access score : how much usable access do I currently have to what the book contains? This is not pages completed. It is closer to “could I retrieve the relevant ideas when I need them?”

For an audit, I would use a deliberately coarse scale:

The labels can stay informal:

What I mean
I have essentially no access to it
I have skimmed it or know its basic territory
I know the central argument or can navigate it
I know it well and can retrieve most of what matters
It is deeply integrated into my working knowledge

Second, a curation weight : how intentionally does this book belong in my present library?

Again, a coarse scale is better than fake precision:

Here means the book clearly supports a current interest, project, practice, or source of delight. A score of means that, if I were honest, the book is just occupying space.

The two scores answer different questions. asks what I can access. asks whether I still have a reason to want access.

Three coordinates for a library

Now split each book’s unit mass into three parts:

and

The identity

holds for every book. Averaging over the whole collection gives three coordinates:

And therefore

is the collection’s realized access: books that belong and that I can already use. is curated potential: books that belong but still sit beyond my present access. is the inactive share: inventory that no longer has a convincing connection to what I care about.

This gives the antilibrary a narrower meaning. It is the curated frontier represented by .

The letter is intentionally less flattering. Call it drift, dead weight, dormant stock, or “books I should stop constructing elaborate excuses for.” The model does not care.

So what should a good library look like?

A good library should probably not maximize any one of these numbers.

Maximizing would give me a collection made entirely of things I already know. Comfortable, perhaps, but not much of an antilibrary. Maximizing would give me a beautifully curated monument to books I never actually use. And maximizing would give me my current storage problem.

The healthier interpretation is dynamic. For a fixed collection, reading a relevant book increases , which moves mass from to without changing . Losing interest in a book decreases , which moves mass towards . Curation removes or replaces that inactive stock, although removing a book also changes , so all three collection-level shares must then be recalculated.

This suggests a few possible theories:

  • The antilibrary as option value. A book in is valuable because it is relevant, available, and waiting for the right question.
  • The antilibrary as a frontier. The balance between and describes the boundary between what I can already use and what I have deliberately placed within reach.
  • The antilibrary as a flow system. Books should move from to as interests become projects, while changes in taste push other books towards . A living library is one in which those movements continue.
  • The antilibrary as an honesty test. Calling every unread book part of an antilibrary is flattering. The weight forces a less convenient question: is this book really intellectual potential, or do I simply dislike admitting that buying it was a mistake?

The model also exposes two opposite failure modes. If is almost zero, the collection may have no unexplored frontier. If is enormous but never becomes , the library may be more aspirational than useful. The interesting quantity may therefore be neither the size of nor the ratio , but the rate at which relevant potential becomes realized access over time.

The next-book problem

There is still a practical question hiding inside all this theory: given the antilibrary I already have, which book should I read next?

For each book, let be my estimate of how useful its ideas would be, and let be the time it would take to read. A very simple priority score is

This favours books that are relevant, still mostly unexplored, potentially useful, and not absurdly expensive in time. The next book would simply be

Of course, this rewards the books whose value I can already see. That creates a problem: I may keep reading variations on things I already understand and never choose the strange book that could change what I care about. To leave some room for curiosity, give each book an uncertainty score and add an exploration bonus:

The parameter measures how adventurous I feel. When , I exploit what already looks useful. As grows, I become more willing to explore books whose value is uncertain. This is a small version of the multi-armed bandit problem: choosing between a promising option and an uncertain one that might teach me something unexpected.

I would not pretend that I can estimate any of these numbers precisely. But even a rough score forces the right questions. Am I avoiding this book because it is irrelevant, because it is long, or because I do not yet understand why it might matter?

Conclusion

The original idea of the antilibrary is powerful because it rescues unread books from the accusation of uselessness. But unreadness alone is too generous. Some unread books are future tools; others are just decor with good intentions.

So perhaps a good library is not one in which everything has been read, nor one in which unread books endlessly accumulate. It is one with a large curated share, a meaningful frontier, and enough movement from potential to realized knowledge to prove that the frontier is alive.

In other words: the goal is not to defeat the antilibrary. It is to make sure it contains the right ignorance.

Further reading

  • Umberto Eco, The Name of the Rose, mostly because it’s a great freaking book.
  • Nassim Nicholas Taleb, The Black Swan, for the argument that unread books can serve as a research tool and a reminder of the limits of our knowledge. And because it’s a great freaking book.
  • My next (future) post, where I am going to analyse my (anti)library and score it using these metrics.