-
Archives
- October 2023
- August 2022
- February 2021
- June 2020
- October 2019
- March 2019
- October 2018
- July 2018
- May 2018
- March 2017
- February 2017
- November 2016
- July 2016
- March 2016
- January 2016
- November 2015
- October 2015
- September 2015
- July 2015
- June 2015
- May 2015
- April 2015
- March 2015
- February 2015
- January 2015
- December 2014
- November 2014
- October 2014
- September 2014
- August 2014
- June 2014
- April 2014
- March 2014
- February 2014
- January 2014
- December 2013
- November 2013
- October 2013
- September 2013
- June 2013
- April 2013
- March 2013
- February 2013
- January 2013
- September 2011
- May 2011
- March 2011
- February 2011
- January 2011
- November 2010
- October 2010
- September 2010
- June 2010
- April 2010
- March 2010
- February 2010
- January 2010
- December 2009
- October 2009
- September 2009
- August 2009
- July 2009
- June 2009
- May 2009
- April 2009
- March 2009
- February 2009
- January 2009
- December 2008
- November 2008
- October 2008
- September 2008
-
Meta
Category Archives: Performance
Beware the cursorMark, my son!
Efficient export of stored text content from multi-shard Solr setups using cursorMark for individual shards and merging the results externally. Continue reading
Posted in eskildsen, Hacking, open source, Performance, Solr
Tagged datasets, Performance, Solr
Leave a comment
DocValues jump tables in Lucene/Solr 8
Lucene/Solr 8 is about to be released. Among a lot of other things is brings LUCENE-8585, written by your truly with a heap of help from Adrien Grand. LUCENE-8585 introduces jump-tables for DocValues, is all about performance and brings speed-ups … Continue reading
Posted in eskildsen, Hacking, Low-level, Lucene, Performance, Solr, Uncategorized
7 Comments
Faster DocValues in Lucene/Solr 7+
This is a fairly technical post explaining LUCENE-8374 and its implications on Lucene, Solr and (qualified guess) Elasticsearch search and retrieval speed. It is primarily relevant for people with indexes of 100M+ documents. Teaser We have a Solr setup for … Continue reading
70TB, 16b docs, 4 machines, 1 SolrCloud
At Statsbiblioteket we maintain a historical net archive for the Danish parts of the Internet. We index it all in Solr and we recently caught up with the present. Time for a status update. The focus is performance and logistics, … Continue reading
Posted in Hacking, Low-level, Performance, Solr, Statsbiblioteket, Uncategorized
6 Comments
The ones that got away
Two and a half ideas of improving Lucene/Solr performance that did not work out. Track the result set bits At the heart of Lucene (and consequently also Solr and ElasticSearch), there is a great amount of doc ID set handling. … Continue reading
Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr
Leave a comment
Speeding up core search
Issue a query, get back the top-X results. It does not get more basic with Solr. So great win if we can improve on that, right? Truth be told, the answer is still “maybe”, but read on for some thoughts, … Continue reading
Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr, Uncategorized
2 Comments
Sampling methods for heuristic faceting
Initial experiments with heuristic faceting in Solr were encouraging: Using just a sample of the result set, it was possible to get correct facet results for large result sets, reducing processing time by an order of magnitude. Alas, further experimentation … Continue reading
Posted in eskildsen, Faceting, Low-level, open source, Performance, Solr
Leave a comment
Dubious guesses, counted correctly
We do have a bit of a performance challenge with heavy faceting on large result sets in our Solr based Net Archive Search. The usual query speed is < 2 seconds, but if the user requests aggregations based on large … Continue reading
Posted in eskildsen, Faceting, Low-level, open source, Performance, Solr
1 Comment
Heuristically correct top-X facets
For most searches in our Net Archive, we have acceptable response time, due to the use of sparse faceting with Solr. Unfortunately as well as expectedly, some of the searches are slow. Response times in minutes slow, if we’re talking … Continue reading
Alternative counter tracking
Warning: Bit-fiddling ahead. The initial driver for implementing Sparse Faceting was to have extraction-time scale with the result set size, instead of with the total number of unique values in the index. From a performance point of view, this works … Continue reading
Posted in eskildsen, Faceting, Hacking, Low-level, Performance, Solr
Leave a comment