Substrate Codebook
A substrate library is a learned dictionary over your work distribution. Whether it compounds is set by the source entropy rate. Zipfian heads compound. Uniform tails do not.
Companion to Substrates: The Dies and Molds of Knowledge Work. Lempel-Ziv intuition applied to operations.
The source-coding bound
A substrate library is a learned dictionary. Each entry compresses a recurring pattern in the work stream into a reusable lookup. The irreducible per-instance work converges to the entropy rate of the source distribution. This is the same bound Lempel-Ziv compression obeys against its source.
Zipfian distributions (few heavy categories carry most mass) compound beautifully. Capture the top categories and the hit rate climbs fast toward a high asymptote. SEV incidents, ticket categories, SKU quality issues are usually like this.
Uniform distributions (many flat categories, no head) do not compound. The library grows without the hit rate climbing. Pure novel research, fully open-ended creative work, sui-generis exceptions are like this.
The operationally useful move is to estimate the entropy rate of your incoming work distribution before investing heavily in a substrate library. If your distribution is Zipfian, the library is a high-leverage investment. If your distribution is flat, you are better off investing in something else.
The factorization escape hatch (the FP / HOF move) is a separate widget topic. When the leaf codebook is sprawling but the entries share parameterized structure, replacing N leaves with one parameterized substrate compresses the codebook from |N| to |g| + |Θ|. Discussed in the parent post.
See also
- Substrates: The Dies and Molds of Knowledge Work - this widget's parent post
- The Knowledge-Work Assembly Line - the framing piece
- The Eighth Ledger: Capitalizing Intellectual Labor - the asset-class story