-
Archives
- October 2023
- August 2022
- February 2021
- June 2020
- October 2019
- March 2019
- October 2018
- July 2018
- May 2018
- March 2017
- February 2017
- November 2016
- July 2016
- March 2016
- January 2016
- November 2015
- October 2015
- September 2015
- July 2015
- June 2015
- May 2015
- April 2015
- March 2015
- February 2015
- January 2015
- December 2014
- November 2014
- October 2014
- September 2014
- August 2014
- June 2014
- April 2014
- March 2014
- February 2014
- January 2014
- December 2013
- November 2013
- October 2013
- September 2013
- June 2013
- April 2013
- March 2013
- February 2013
- January 2013
- September 2011
- May 2011
- March 2011
- February 2011
- January 2011
- November 2010
- October 2010
- September 2010
- June 2010
- April 2010
- March 2010
- February 2010
- January 2010
- December 2009
- October 2009
- September 2009
- August 2009
- July 2009
- June 2009
- May 2009
- April 2009
- March 2009
- February 2009
- January 2009
- December 2008
- November 2008
- October 2008
- September 2008
-
Meta
Category Archives: open source
Beware the cursorMark, my son!
Efficient export of stored text content from multi-shard Solr setups using cursorMark for individual shards and merging the results externally. Continue reading
Posted in eskildsen, Hacking, open source, Performance, Solr
Tagged datasets, Performance, Solr
Leave a comment
The ones that got away
Two and a half ideas of improving Lucene/Solr performance that did not work out. Track the result set bits At the heart of Lucene (and consequently also Solr and ElasticSearch), there is a great amount of doc ID set handling. … Continue reading
Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr
Leave a comment
Speeding up core search
Issue a query, get back the top-X results. It does not get more basic with Solr. So great win if we can improve on that, right? Truth be told, the answer is still “maybe”, but read on for some thoughts, … Continue reading
Posted in eskildsen, Hacking, Low-level, Lucene, open source, Performance, Solr, Uncategorized
2 Comments
Sampling methods for heuristic faceting
Initial experiments with heuristic faceting in Solr were encouraging: Using just a sample of the result set, it was possible to get correct facet results for large result sets, reducing processing time by an order of magnitude. Alas, further experimentation … Continue reading
Posted in eskildsen, Faceting, Low-level, open source, Performance, Solr
Leave a comment
Dubious guesses, counted correctly
We do have a bit of a performance challenge with heavy faceting on large result sets in our Solr based Net Archive Search. The usual query speed is < 2 seconds, but if the user requests aggregations based on large … Continue reading
Posted in eskildsen, Faceting, Low-level, open source, Performance, Solr
1 Comment
Net Archive Search building blocks
An extremely webarchive-discovery and Statsbiblioteket centric description of some of the technical possibilities with Net Archive Search. This could be considered internal documentation, but we like to share. There are currently 2 generations of indexes at Statsbiblioteket: v1 (22TB) & … Continue reading
Posted in eskildsen, open source, Solr
2 Comments
Sparse facet caching
As explained in Ten times faster, distributed faceting in standard Solr is two-phase: Each shard performs standard faceting and returns the top limit*1.5+10 terms. The merger calculates the top limit terms. Standard faceting is a two-step process: For each term … Continue reading
Posted in eskildsen, Faceting, Hacking, Low-level, open source, Performance, Solr
3 Comments
Ten times faster
One week ago I complained about Solr’s two-phase distributed faceting being slow in the second phase – ten times slower than the first phase. The culprit was the fine-counting of top-X terms, with each term-count being done as an intersection … Continue reading
Posted in eskildsen, Faceting, Hacking, Low-level, open source, Performance, Solr, Uncategorized
5 Comments
Sparse facet counting without the downsides
The SOLR-5894 issue “Speed up high-cardinality facets with sparse counters” is close to being functionally complete (facet.method=fcs and facet.sort=index still pending). This post explains the different tricks used in the implementation and their impact on performance. Baseline Most of the … Continue reading
Posted in eskildsen, Faceting, Low-level, open source, Performance, Solr
2 Comments