https://bugs.koha-community.org/bugzilla3/show_bug.cgi?id=42835 --- Comment #63 from Luis Bataller <luis.bataller@xercode.es> --- Hi, Thanks for picking this up — indexing item data in Elasticsearch is something Koha has needed for a long time. We went through this same design question some months ago and evaluated both options: a separate item index, and items as `nested` objects inside the biblio document. We ended up choosing `nested`. We picked up the existing bug 8589 to offer that work to the community, but we never attached the patches there — we offered them and then let it drop, mostly on our side. The implementation is now in pre-production on an installation with roughly 1.5 million biblio records and 8 million items. A search filtered only by item attributes works fine with a separate index — that matches what bug 42835 currently implements, as far as I can tell. The difficulty we hit was with queries mixing item-level and biblio-level criteria, which is the common shape in both the OPAC and the staff interface. What follows is a summary of the analysis we carried out against that catalogue. --- Our reference case was the most frequent one: **library = b, title like "..."**. In that catalogue the largest library holds around half a million items, which resolves to somewhere between 140,000 and 320,000 distinct biblionumbers depending on the local items-per-record ratio — 10-20% of the whole catalogue in a single intermediate set. Two problems followed: 1. **Intermediate set size.** Passing that many biblionumbers as a `terms` filter exceeds `index.max_terms_count` (65,536 by default). The workarounds we considered were splitting the query into chunks, which breaks global scoring and relevance ordering, or a `terms lookup`, which introduces a synchronous write plus a refresh in the middle of a read path. 2. **Relevance pagination.** The selective criterion is the title, so the biblio query has to run first, but then there is no way to know how many of its hits will survive the item filter. Filling a page of 20 results becomes an iterative loop with an unpredictable number of round-trips, or requires fetching the full candidate set and re-sorting in application code, losing BM25 scoring. --- So my main question is: have you already worked out how general searches combining item and biblio criteria would be resolved, particularly these two points? If you have a solution, I would genuinely like to understand it — the separate-index approach has real advantages we would benefit from, particularly a much lower cost for full reindexing when an item field is added. If it turns out to be an open problem, we would be glad to share what we have. The patches are not attached to bug 8589 yet, but we can post them there promptly if there is interest, whether as something to contribute or just as a reference for comparison. We are also happy to run comparative tests against our dataset if that would help validate either design. Thanks again for working on this. -- You are receiving this mail because: You are watching all bug changes.