Category Archives: Hacking

Beware the cursorMark, my son!

Efficient export of stored text content from multi-shard Solr setups using cursorMark for individual shards and merging the results externally. Continue reading

Posted in eskildsen, Hacking, open source, Performance, Solr | Tagged , , | Leave a comment

Dumb-down at Indexing or Nested Data in the Solr Search Engine

Sigfrid Lundberg, Ph. D., Software Developer Royal Danish Library Copenhagen Denmark twitter — github — web site Are passions, then, the Pagans of the soul? Reason alone baptized? alone ordain’d To touch things sacred? (Edward Young — 1683-1765) Introduction The … Continue reading

Posted in sigge, Solr, usability | Leave a comment

SolrWayback 4.0 release! What’s it all about? Part 2

In this blog post I will go into the more technical details of SolrWayback and the new version 4.0 release. The whole frontend GUI was rewritten from scratch to be up to date with 2020 web-applications expectations along with many … Continue reading

Posted in open source, Solr, Visualization, Web | Tagged , , , | 1 Comment

Which type bug?

A light tale of bug hunting an Out Of Memory problem with SolrCloud. The setup and the problem At the Royal Danish Library we provide full text search for the Danish Netarchive. The heavy lifting is done in a single … Continue reading

Posted in eskildsen, Solr | Tagged , | Leave a comment

DocValues jump tables in Lucene/Solr 8

Lucene/Solr 8 is about to be released. Among a lot of other things is brings LUCENE-8585, written by your truly with a heap of help from Adrien Grand. LUCENE-8585 introduces jump-tables for DocValues, is all about performance and brings speed-ups … Continue reading

Posted in eskildsen, Hacking, Low-level, Lucene, Performance, Solr, Uncategorized | 7 Comments

Faster DocValues in Lucene/Solr 7+

This is a fairly technical post explaining LUCENE-8374 and its implications on Lucene, Solr and (qualified guess) Elasticsearch search and retrieval speed. It is primarily relevant for people with indexes of 100M+ documents. Teaser We have a Solr setup for … Continue reading

Posted in eskildsen, Hacking, Low-level, Lucene, Performance, Solr | 1 Comment

Visualising Netarchive Harvests

  An overview of website harvest data is important for both research and development operations in the netarchive team at Det Kgl. Bibliotek. In this post we present a recent frontend visualisation widget we have made. From the SolrWayback Machine … Continue reading

Posted in Blogging, Solr, Web | Tagged , | Leave a comment

70TB, 16b docs, 4 machines, 1 SolrCloud

At Statsbiblioteket we maintain a historical net archive for the Danish parts of the Internet. We index it all in Solr and we recently caught up with the present. Time for a status update. The focus is performance and logistics, … Continue reading

Posted in Hacking, Low-level, Performance, Solr, Statsbiblioteket, Uncategorized | 6 Comments

The ones that got away

Two and a half ideas of improving Lucene/Solr performance that did not work out. Track the result set bits At the heart of Lucene (and consequently also Solr and ElasticSearch), there is a great amount of doc ID set handling. … Continue reading

Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr | Leave a comment

Speeding up core search

Issue a query, get back the top-X results. It does not get more basic with Solr. So great win if we can improve on that, right? Truth be told, the answer is still “maybe”, but read on for some thoughts, … Continue reading

Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr, Uncategorized | 2 Comments