refactor: native fork-safe threads#14163
Conversation
|
|
7b63a8c to
8f151dd
Compare
Performance SLOsComparing candidate refactor/native-forksafe-threads (881c153) with baseline main (3cc3eb7) 📈 Performance Regressions (2 suites)📈 iastaspects - 118/118✅ add_aspectTime: ✅ 103.017µs (SLO: <130.000µs 📉 -20.8%) vs baseline: -2.2% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +3.9% ✅ add_inplace_aspectTime: ✅ 100.841µs (SLO: <130.000µs 📉 -22.4%) vs baseline: -4.2% Memory: ✅ 42.861MB (SLO: <46.000MB -6.8%) vs baseline: +4.3% ✅ add_inplace_noaspectTime: ✅ 28.382µs (SLO: <40.000µs 📉 -29.0%) vs baseline: +0.2% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ add_noaspectTime: ✅ 48.682µs (SLO: <70.000µs 📉 -30.5%) vs baseline: -1.2% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ bytearray_aspectTime: ✅ 251.087µs (SLO: <400.000µs 📉 -37.2%) vs baseline: +0.7% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ bytearray_extend_aspectTime: ✅ 635.720µs (SLO: <800.000µs 📉 -20.5%) vs baseline: -0.3% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ bytearray_extend_noaspectTime: ✅ 266.082µs (SLO: <400.000µs 📉 -33.5%) vs baseline: +0.5% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ bytearray_noaspectTime: ✅ 134.869µs (SLO: <300.000µs 📉 -55.0%) vs baseline: ~same Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ bytes_aspectTime: ✅ 218.231µs (SLO: <300.000µs 📉 -27.3%) vs baseline: +0.8% Memory: ✅ 42.841MB (SLO: <46.000MB -6.9%) vs baseline: +4.2% ✅ bytes_noaspectTime: ✅ 133.334µs (SLO: <200.000µs 📉 -33.3%) vs baseline: +1.0% Memory: ✅ 42.841MB (SLO: <46.000MB -6.9%) vs baseline: +4.1% ✅ bytesio_aspectTime: ✅ 3.774ms (SLO: <5.000ms 📉 -24.5%) vs baseline: -0.9% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.4% ✅ bytesio_noaspectTime: ✅ 315.126µs (SLO: <420.000µs 📉 -25.0%) vs baseline: +0.1% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.0% ✅ capitalize_aspectTime: ✅ 89.561µs (SLO: <300.000µs 📉 -70.1%) vs baseline: +1.5% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.1% ✅ capitalize_noaspectTime: ✅ 252.176µs (SLO: <300.000µs 📉 -15.9%) vs baseline: +0.3% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ casefold_aspectTime: ✅ 89.016µs (SLO: <500.000µs 📉 -82.2%) vs baseline: +0.6% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.3% ✅ casefold_noaspectTime: ✅ 306.523µs (SLO: <500.000µs 📉 -38.7%) vs baseline: -0.1% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ decode_aspectTime: ✅ 86.931µs (SLO: <100.000µs 📉 -13.1%) vs baseline: ~same Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.3% ✅ decode_noaspectTime: ✅ 152.010µs (SLO: <210.000µs 📉 -27.6%) vs baseline: +0.2% Memory: ✅ 42.861MB (SLO: <46.000MB -6.8%) vs baseline: +4.1% ✅ encode_aspectTime: ✅ 84.630µs (SLO: <200.000µs 📉 -57.7%) vs baseline: -0.4% Memory: ✅ 42.841MB (SLO: <46.000MB -6.9%) vs baseline: +3.9% ✅ encode_noaspectTime: ✅ 140.519µs (SLO: <200.000µs 📉 -29.7%) vs baseline: +2.2% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.1% ✅ format_aspectTime: ✅ 14.613ms (SLO: <19.200ms 📉 -23.9%) vs baseline: +0.1% Memory: ✅ 43.175MB (SLO: <46.000MB -6.1%) vs baseline: +4.7% ✅ format_map_aspectTime: ✅ 16.416ms (SLO: <21.500ms 📉 -23.6%) vs baseline: +0.4% Memory: ✅ 42.979MB (SLO: <46.000MB -6.6%) vs baseline: +4.3% ✅ format_map_noaspectTime: ✅ 374.539µs (SLO: <500.000µs 📉 -25.1%) vs baseline: -0.2% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ format_noaspectTime: ✅ 302.943µs (SLO: <500.000µs 📉 -39.4%) vs baseline: -0.5% Memory: ✅ 42.802MB (SLO: <46.000MB -7.0%) vs baseline: +4.3% ✅ index_aspectTime: ✅ 123.744µs (SLO: <300.000µs 📉 -58.8%) vs baseline: +1.2% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.2% ✅ index_noaspectTime: ✅ 40.420µs (SLO: <300.000µs 📉 -86.5%) vs baseline: +1.0% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.5% ✅ join_aspectTime: ✅ 212.427µs (SLO: <300.000µs 📉 -29.2%) vs baseline: +0.3% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.6% ✅ join_noaspectTime: ✅ 140.596µs (SLO: <300.000µs 📉 -53.1%) vs baseline: -1.6% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ ljust_aspectTime: ✅ 577.954µs (SLO: <700.000µs 📉 -17.4%) vs baseline: 📈 +16.5% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.0% ✅ ljust_noaspectTime: ✅ 258.335µs (SLO: <300.000µs 📉 -13.9%) vs baseline: +0.9% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.4% ✅ lower_aspectTime: ✅ 293.979µs (SLO: <500.000µs 📉 -41.2%) vs baseline: +0.3% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.4% ✅ lower_noaspectTime: ✅ 236.210µs (SLO: <300.000µs 📉 -21.3%) vs baseline: +0.6% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.4% ✅ lstrip_aspectTime: ✅ 0.272ms (SLO: <3.000ms 📉 -90.9%) vs baseline: +0.9% Memory: ✅ 42.841MB (SLO: <46.000MB -6.9%) vs baseline: +4.0% ✅ lstrip_noaspectTime: ✅ 0.178ms (SLO: <3.000ms 📉 -94.1%) vs baseline: +1.6% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ modulo_aspectTime: ✅ 14.363ms (SLO: <18.750ms 📉 -23.4%) vs baseline: +0.7% Memory: ✅ 42.979MB (SLO: <46.000MB -6.6%) vs baseline: +4.2% ✅ modulo_aspect_for_bytearray_bytearrayTime: ✅ 14.778ms (SLO: <19.350ms 📉 -23.6%) vs baseline: +0.3% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.2% ✅ modulo_aspect_for_bytesTime: ✅ 14.390ms (SLO: <18.900ms 📉 -23.9%) vs baseline: -0.2% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.2% ✅ modulo_aspect_for_bytes_bytearrayTime: ✅ 14.571ms (SLO: <19.150ms 📉 -23.9%) vs baseline: +0.3% Memory: ✅ 42.998MB (SLO: <46.000MB -6.5%) vs baseline: +4.5% ✅ modulo_noaspectTime: ✅ 0.360ms (SLO: <3.000ms 📉 -88.0%) vs baseline: -1.6% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.3% ✅ replace_aspectTime: ✅ 18.398ms (SLO: <24.000ms 📉 -23.3%) vs baseline: -0.3% Memory: ✅ 42.979MB (SLO: <46.000MB -6.6%) vs baseline: +4.4% ✅ replace_noaspectTime: ✅ 281.567µs (SLO: <300.000µs -6.1%) vs baseline: +1.2% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.5% ✅ repr_aspectTime: ✅ 312.325µs (SLO: <420.000µs 📉 -25.6%) vs baseline: +0.6% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ repr_noaspectTime: ✅ 47.073µs (SLO: <90.000µs 📉 -47.7%) vs baseline: -0.2% Memory: ✅ 42.979MB (SLO: <46.000MB -6.6%) vs baseline: +4.6% ✅ rstrip_aspectTime: ✅ 383.545µs (SLO: <500.000µs 📉 -23.3%) vs baseline: +0.5% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.1% ✅ rstrip_noaspectTime: ✅ 184.303µs (SLO: <300.000µs 📉 -38.6%) vs baseline: +2.3% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ slice_aspectTime: ✅ 185.036µs (SLO: <300.000µs 📉 -38.3%) vs baseline: +1.4% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.4% ✅ slice_noaspectTime: ✅ 54.491µs (SLO: <90.000µs 📉 -39.5%) vs baseline: +1.0% Memory: ✅ 42.939MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ stringio_aspectTime: ✅ 4.407ms (SLO: <5.000ms 📉 -11.9%) vs baseline: 📈 +14.4% Memory: ✅ 42.979MB (SLO: <46.000MB -6.6%) vs baseline: +4.4% ✅ stringio_noaspectTime: ✅ 346.454µs (SLO: <500.000µs 📉 -30.7%) vs baseline: -0.5% Memory: ✅ 42.821MB (SLO: <46.000MB -6.9%) vs baseline: +4.0% ✅ strip_aspectTime: ✅ 269.112µs (SLO: <350.000µs 📉 -23.1%) vs baseline: -1.1% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.5% ✅ strip_noaspectTime: ✅ 176.534µs (SLO: <240.000µs 📉 -26.4%) vs baseline: +1.2% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.2% ✅ swapcase_aspectTime: ✅ 333.976µs (SLO: <500.000µs 📉 -33.2%) vs baseline: +0.5% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.4% ✅ swapcase_noaspectTime: ✅ 270.936µs (SLO: <400.000µs 📉 -32.3%) vs baseline: -0.4% Memory: ✅ 42.998MB (SLO: <46.000MB -6.5%) vs baseline: +4.5% ✅ title_aspectTime: ✅ 319.943µs (SLO: <500.000µs 📉 -36.0%) vs baseline: -0.1% Memory: ✅ 42.900MB (SLO: <46.000MB -6.7%) vs baseline: +4.3% ✅ title_noaspectTime: ✅ 256.967µs (SLO: <400.000µs 📉 -35.8%) vs baseline: -2.2% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.1% ✅ translate_aspectTime: ✅ 492.924µs (SLO: <700.000µs 📉 -29.6%) vs baseline: +0.4% Memory: ✅ 42.861MB (SLO: <46.000MB -6.8%) vs baseline: +4.2% ✅ translate_noaspectTime: ✅ 424.654µs (SLO: <500.000µs 📉 -15.1%) vs baseline: -0.5% Memory: ✅ 42.920MB (SLO: <46.000MB -6.7%) vs baseline: +4.4% ✅ upper_aspectTime: ✅ 294.840µs (SLO: <500.000µs 📉 -41.0%) vs baseline: +0.6% Memory: ✅ 42.959MB (SLO: <46.000MB -6.6%) vs baseline: +4.5% ✅ upper_noaspectTime: ✅ 235.633µs (SLO: <400.000µs 📉 -41.1%) vs baseline: +0.6% Memory: ✅ 42.880MB (SLO: <46.000MB -6.8%) vs baseline: +4.3% 📈 iastaspectsospath - 24/24✅ ospathbasename_aspectTime: ✅ 507.776µs (SLO: <700.000µs 📉 -27.5%) vs baseline: 📈 +21.7% Memory: ✅ 42.585MB (SLO: <46.000MB -7.4%) vs baseline: +4.1% ✅ ospathbasename_noaspectTime: ✅ 432.158µs (SLO: <700.000µs 📉 -38.3%) vs baseline: +0.7% Memory: ✅ 42.644MB (SLO: <46.000MB -7.3%) vs baseline: +4.7% ✅ ospathjoin_aspectTime: ✅ 627.353µs (SLO: <700.000µs 📉 -10.4%) vs baseline: +0.9% Memory: ✅ 42.566MB (SLO: <46.000MB -7.5%) vs baseline: +4.3% ✅ ospathjoin_noaspectTime: ✅ 631.404µs (SLO: <700.000µs -9.8%) vs baseline: -0.3% Memory: ✅ 42.566MB (SLO: <46.000MB -7.5%) vs baseline: +4.4% ✅ ospathnormcase_aspectTime: ✅ 348.248µs (SLO: <700.000µs 📉 -50.3%) vs baseline: +1.2% Memory: ✅ 42.507MB (SLO: <46.000MB -7.6%) vs baseline: +4.3% ✅ ospathnormcase_noaspectTime: ✅ 357.185µs (SLO: <700.000µs 📉 -49.0%) vs baseline: +0.6% Memory: ✅ 42.723MB (SLO: <46.000MB -7.1%) vs baseline: +4.9% ✅ ospathsplit_aspectTime: ✅ 486.752µs (SLO: <700.000µs 📉 -30.5%) vs baseline: +1.2% Memory: ✅ 42.546MB (SLO: <46.000MB -7.5%) vs baseline: +4.5% ✅ ospathsplit_noaspectTime: ✅ 498.017µs (SLO: <700.000µs 📉 -28.9%) vs baseline: +1.5% Memory: ✅ 42.585MB (SLO: <46.000MB -7.4%) vs baseline: +4.1% ✅ ospathsplitdrive_aspectTime: ✅ 376.190µs (SLO: <700.000µs 📉 -46.3%) vs baseline: +1.5% Memory: ✅ 42.566MB (SLO: <46.000MB -7.5%) vs baseline: +4.5% ✅ ospathsplitdrive_noaspectTime: ✅ 72.831µs (SLO: <700.000µs 📉 -89.6%) vs baseline: -0.4% Memory: ✅ 42.546MB (SLO: <46.000MB -7.5%) vs baseline: +4.2% ✅ ospathsplitext_aspectTime: ✅ 456.264µs (SLO: <700.000µs 📉 -34.8%) vs baseline: +2.0% Memory: ✅ 42.605MB (SLO: <46.000MB -7.4%) vs baseline: +4.6% ✅ ospathsplitext_noaspectTime: ✅ 463.925µs (SLO: <700.000µs 📉 -33.7%) vs baseline: +1.3% Memory: ✅ 42.605MB (SLO: <46.000MB -7.4%) vs baseline: +4.2% 🟡 Near SLO Breach (1 suite)🟡 tracer - 6/6✅ largeTime: ✅ 31.786ms (SLO: <32.950ms -3.5%) vs baseline: +0.6% Memory: ✅ 36.766MB (SLO: <39.250MB -6.3%) vs baseline: +4.4% ✅ mediumTime: ✅ 3.129ms (SLO: <3.200ms -2.2%) vs baseline: ~same Memory: ✅ 35.547MB (SLO: <38.750MB -8.3%) vs baseline: +4.4% ✅ smallTime: ✅ 363.966µs (SLO: <370.000µs 🟡 -1.6%) vs baseline: +3.1% Memory: ✅ 35.488MB (SLO: <38.750MB -8.4%) vs baseline: +4.5% 📉 Performance Improvements (2 suites)📉 httppropagationextract - 60/60✅ all_styles_all_headersTime: ✅ 79.572µs (SLO: <100.000µs 📉 -20.4%) vs baseline: -1.3% Memory: ✅ 35.625MB (SLO: <38.000MB -6.2%) vs baseline: +4.2% ✅ b3_headersTime: ✅ 12.770µs (SLO: <20.000µs 📉 -36.2%) vs baseline: 📉 -10.4% Memory: ✅ 35.665MB (SLO: <38.000MB -6.1%) vs baseline: +4.4% ✅ b3_single_headersTime: ✅ 11.819µs (SLO: <20.000µs 📉 -40.9%) vs baseline: 📉 -11.0% Memory: ✅ 35.527MB (SLO: <38.000MB -6.5%) vs baseline: +4.0% ✅ datadog_tracecontext_tracestate_not_propagated_on_trace_id_no_matchTime: ✅ 60.705µs (SLO: <80.000µs 📉 -24.1%) vs baseline: -5.4% Memory: ✅ 35.606MB (SLO: <38.000MB -6.3%) vs baseline: +4.2% ✅ datadog_tracecontext_tracestate_propagated_on_trace_id_matchTime: ✅ 62.553µs (SLO: <80.000µs 📉 -21.8%) vs baseline: -5.3% Memory: ✅ 35.566MB (SLO: <38.000MB -6.4%) vs baseline: +4.2% ✅ empty_headersTime: ✅ 1.304µs (SLO: <10.000µs 📉 -87.0%) vs baseline: 📉 -17.5% Memory: ✅ 35.625MB (SLO: <38.000MB -6.2%) vs baseline: +4.5% ✅ full_t_id_datadog_headersTime: ✅ 20.848µs (SLO: <30.000µs 📉 -30.5%) vs baseline: -8.0% Memory: ✅ 35.645MB (SLO: <38.000MB -6.2%) vs baseline: +4.5% ✅ invalid_priority_headerTime: ✅ 5.928µs (SLO: <10.000µs 📉 -40.7%) vs baseline: -8.9% Memory: ✅ 35.684MB (SLO: <38.000MB -6.1%) vs baseline: +4.7% ✅ invalid_span_id_headerTime: ✅ 5.873µs (SLO: <10.000µs 📉 -41.3%) vs baseline: -9.7% Memory: ✅ 35.606MB (SLO: <38.000MB -6.3%) vs baseline: +4.5% ✅ invalid_tags_headerTime: ✅ 5.869µs (SLO: <10.000µs 📉 -41.3%) vs baseline: -10.0% Memory: ✅ 35.684MB (SLO: <38.000MB -6.1%) vs baseline: +4.5% ✅ invalid_trace_id_headerTime: ✅ 5.892µs (SLO: <10.000µs 📉 -41.1%) vs baseline: 📉 -10.2% Memory: ✅ 35.724MB (SLO: <38.000MB -6.0%) vs baseline: +4.7% ✅ large_header_no_matchesTime: ✅ 27.001µs (SLO: <30.000µs -10.0%) vs baseline: -2.4% Memory: ✅ 35.704MB (SLO: <38.000MB -6.0%) vs baseline: +4.5% ✅ large_valid_headers_allTime: ✅ 28.240µs (SLO: <40.000µs 📉 -29.4%) vs baseline: -1.6% Memory: ✅ 35.743MB (SLO: <38.000MB -5.9%) vs baseline: +4.7% ✅ medium_header_no_matchesTime: ✅ 9.204µs (SLO: <20.000µs 📉 -54.0%) vs baseline: -6.4% Memory: ✅ 35.724MB (SLO: <38.000MB -6.0%) vs baseline: +4.5% ✅ medium_valid_headers_allTime: ✅ 10.695µs (SLO: <20.000µs 📉 -46.5%) vs baseline: -5.3% Memory: ✅ 35.586MB (SLO: <38.000MB -6.4%) vs baseline: +4.1% ✅ none_propagation_styleTime: ✅ 1.406µs (SLO: <10.000µs 📉 -85.9%) vs baseline: 📉 -17.2% Memory: ✅ 35.665MB (SLO: <38.000MB -6.1%) vs baseline: +4.6% ✅ tracecontext_headersTime: ✅ 32.573µs (SLO: <40.000µs 📉 -18.6%) vs baseline: -6.0% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.2% ✅ valid_headers_allTime: ✅ 5.875µs (SLO: <10.000µs 📉 -41.2%) vs baseline: -9.3% Memory: ✅ 35.665MB (SLO: <38.000MB -6.1%) vs baseline: +4.5% ✅ valid_headers_basicTime: ✅ 5.473µs (SLO: <10.000µs 📉 -45.3%) vs baseline: -10.0% Memory: ✅ 35.527MB (SLO: <38.000MB -6.5%) vs baseline: +4.1% ✅ wsgi_empty_headersTime: ✅ 1.305µs (SLO: <10.000µs 📉 -87.0%) vs baseline: 📉 -18.0% Memory: ✅ 35.724MB (SLO: <38.000MB -6.0%) vs baseline: +4.7% ✅ wsgi_invalid_priority_headerTime: ✅ 5.933µs (SLO: <10.000µs 📉 -40.7%) vs baseline: -9.6% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.3% ✅ wsgi_invalid_span_id_headerTime: ✅ 1.298µs (SLO: <10.000µs 📉 -87.0%) vs baseline: 📉 -18.9% Memory: ✅ 35.645MB (SLO: <38.000MB -6.2%) vs baseline: +4.2% ✅ wsgi_invalid_tags_headerTime: ✅ 5.941µs (SLO: <10.000µs 📉 -40.6%) vs baseline: -9.7% Memory: ✅ 35.625MB (SLO: <38.000MB -6.2%) vs baseline: +4.3% ✅ wsgi_invalid_trace_id_headerTime: ✅ 5.945µs (SLO: <10.000µs 📉 -40.5%) vs baseline: -9.4% Memory: ✅ 35.606MB (SLO: <38.000MB -6.3%) vs baseline: +4.2% ✅ wsgi_large_header_no_matchesTime: ✅ 28.034µs (SLO: <40.000µs 📉 -29.9%) vs baseline: -2.3% Memory: ✅ 35.606MB (SLO: <38.000MB -6.3%) vs baseline: +4.3% ✅ wsgi_large_valid_headers_allTime: ✅ 29.131µs (SLO: <40.000µs 📉 -27.2%) vs baseline: -2.7% Memory: ✅ 35.665MB (SLO: <38.000MB -6.1%) vs baseline: +4.5% ✅ wsgi_medium_header_no_matchesTime: ✅ 9.418µs (SLO: <20.000µs 📉 -52.9%) vs baseline: -7.4% Memory: ✅ 35.645MB (SLO: <38.000MB -6.2%) vs baseline: +4.6% ✅ wsgi_medium_valid_headers_allTime: ✅ 11.048µs (SLO: <20.000µs 📉 -44.8%) vs baseline: -4.1% Memory: ✅ 35.566MB (SLO: <38.000MB -6.4%) vs baseline: +4.1% ✅ wsgi_valid_headers_allTime: ✅ 5.938µs (SLO: <10.000µs 📉 -40.6%) vs baseline: -9.4% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +3.9% ✅ wsgi_valid_headers_basicTime: ✅ 5.513µs (SLO: <10.000µs 📉 -44.9%) vs baseline: -9.6% Memory: ✅ 35.566MB (SLO: <38.000MB -6.4%) vs baseline: +4.1% 📉 telemetryaddmetric - 30/30✅ 1-count-metric-1-timesTime: ✅ 2.258µs (SLO: <20.000µs 📉 -88.7%) vs baseline: 📉 -23.6% Memory: ✅ 35.566MB (SLO: <38.000MB -6.4%) vs baseline: +4.3% ✅ 1-count-metrics-100-timesTime: ✅ 149.006µs (SLO: <220.000µs 📉 -32.3%) vs baseline: 📉 -27.1% Memory: ✅ 35.468MB (SLO: <38.000MB -6.7%) vs baseline: +3.9% ✅ 1-distribution-metric-1-timesTime: ✅ 2.472µs (SLO: <20.000µs 📉 -87.6%) vs baseline: 📉 -24.9% Memory: ✅ 35.586MB (SLO: <38.000MB -6.4%) vs baseline: +4.5% ✅ 1-distribution-metrics-100-timesTime: ✅ 164.557µs (SLO: <230.000µs 📉 -28.5%) vs baseline: 📉 -23.1% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.1% ✅ 1-gauge-metric-1-timesTime: ✅ 1.956µs (SLO: <20.000µs 📉 -90.2%) vs baseline: 📉 -11.0% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.3% ✅ 1-gauge-metrics-100-timesTime: ✅ 135.484µs (SLO: <150.000µs -9.7%) vs baseline: -1.7% Memory: ✅ 35.527MB (SLO: <38.000MB -6.5%) vs baseline: +4.3% ✅ 1-rate-metric-1-timesTime: ✅ 2.222µs (SLO: <20.000µs 📉 -88.9%) vs baseline: 📉 -27.8% Memory: ✅ 35.586MB (SLO: <38.000MB -6.4%) vs baseline: +4.6% ✅ 1-rate-metrics-100-timesTime: ✅ 161.951µs (SLO: <250.000µs 📉 -35.2%) vs baseline: 📉 -25.6% Memory: ✅ 35.507MB (SLO: <38.000MB -6.6%) vs baseline: +4.1% ✅ 100-count-metrics-100-timesTime: ✅ 15.193ms (SLO: <22.000ms 📉 -30.9%) vs baseline: 📉 -25.3% Memory: ✅ 35.704MB (SLO: <38.000MB -6.0%) vs baseline: +4.7% ✅ 100-distribution-metrics-100-timesTime: ✅ 1.748ms (SLO: <2.550ms 📉 -31.5%) vs baseline: 📉 -22.6% Memory: ✅ 35.665MB (SLO: <38.000MB -6.1%) vs baseline: +4.7% ✅ 100-gauge-metrics-100-timesTime: ✅ 1.402ms (SLO: <1.550ms -9.6%) vs baseline: -0.3% Memory: ✅ 35.625MB (SLO: <38.000MB -6.2%) vs baseline: +4.9% ✅ 100-rate-metrics-100-timesTime: ✅ 1.706ms (SLO: <2.550ms 📉 -33.1%) vs baseline: 📉 -23.3% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.6% ✅ flush-1-metricTime: ✅ 3.582µs (SLO: <20.000µs 📉 -82.1%) vs baseline: 📉 -18.7% Memory: ✅ 35.547MB (SLO: <38.000MB -6.5%) vs baseline: +4.1% ✅ flush-100-metricsTime: ✅ 174.506µs (SLO: <250.000µs 📉 -30.2%) vs baseline: ~same Memory: ✅ 35.586MB (SLO: <38.000MB -6.4%) vs baseline: +4.5% ✅ flush-1000-metricsTime: ✅ 2.192ms (SLO: <2.500ms 📉 -12.3%) vs baseline: +0.4% Memory: ✅ 36.392MB (SLO: <38.750MB -6.1%) vs baseline: +4.6%
|
8f151dd to
ced3c89
Compare
d00259c to
9967a4d
Compare
|
One thing that came up in our recent discussion I want to re-state here: I think we need to make sure that starting and stopping the periodic threads doesn't indefinitely delay the periodic callbacks. Specifically, thinking of something like profiling, where we upload data every 60 seconds. If we exit the thread, then the 60 second timer is restarted when the thread is restarted. And so if a program forks more than once every 60 seconds we will never actually upload a profile from the parent process. I think if we stick to this approach, we probably want a deadline for the periodic thread. And the time the thread sleeps is relative to the time left until the deadline. The deadline wouldn't reset by default when stopping/restarting the thread, only when the callback is called. LMK if that makes sense, I'm happy to sketch it out more. |
does 817730d8796cf406651590daf9a2ffd94050bbf9 |
We refactor the native periodict thread implementation to be fork-safe, the sense that all such threads that are running at the time of a fork are automatically stopped before the fork, then restarted after it. We also take care to avoid stopping and restarting threads in the parent process if we detect an immediate call to fork again.
That was fast! From my tests, the accumulation is fixed, and the average overhead lowered to only 13% |
taegyunkim
left a comment
There was a problem hiding this comment.
Sorry that I'm late! A few minor points.
|
Haven't had a chance to look at it yet, but we got a coredump from https://gitlab.ddbuild.io/DataDog/apm-reliability/dd-trace-py/-/jobs/1412900511 this debugger suite |
That seems to be unrelated to this change, but let's save the dump so that we can investigate |
6aa6235 to
8501928
Compare
8501928 to
881c153
Compare
…er fork() (#16545) ## Description After #14163, we started to see a profiler behavior change in `dd-trace-doe` CI: - `TestLanguage/python/flush/false` expected `0` profiles, got `3`. https://github.com/DataDog/dd-trace-doe/actions/runs/22102391228/job/63875669025 - During pre-fork, `PeriodicThread.stop()` sets the internal request event to wake the periodic threads promptly. - After fork restart, that same request could be interpreted as a real wakeup, causing an immediate `periodic()` run. - For profiling, that means an immediate `upload()` right after fork, which results in unexpected profile exports. What this PR changfes: - Distinguishes a fork-stop wakeup from a real `awake()` request in `PeriodicThread` - Clears the request after fork only when it came from the pre-fork stop path. - Preserves real `awake()` behavior across restart windows. ## Testing <!-- Describe your testing strategy or note what tests are included --> ## Risks <!-- Note any risks associated with this change, or "None" if no risks --> ## Additional Notes <!-- Any other information that would be helpful for reviewers --> Co-authored-by: taegyun.kim <[email protected]>
During Python finalization, CPython's take_gil calls pthread_exit() on non-main threads attempting to reacquire the GIL. On glibc, pthread_exit is implemented via abi::__forced_unwind — a forced stack unwind that, if it escapes a std::thread callable without proper handling, triggers std::terminate and a SIGABRT. This was made reachable by the native fork-safe threads refactor (PR #14163) which actively restarts periodic threads in forked child processes. When such a child exits, threads may be mid-callback when finalization begins, and CPython kills them via pthread_exit before the atexit handler's stop signal can be processed. Changes: - Catch abi::__forced_unwind in the thread lambda (glibc only), signal _stopped so join() unblocks, and re-throw to let the unwind complete through libstdc++'s std::thread wrapper. - Make _atexit() tolerate _thread == nullptr instead of raising. - Check start() return value in _after_fork() to avoid silent failures. - Add try/except in the Python atexit handler so one thread's failure does not skip stopping the rest.
…er fork() (#16545) ## Description After #14163, we started to see a profiler behavior change in `dd-trace-doe` CI: - `TestLanguage/python/flush/false` expected `0` profiles, got `3`. https://github.com/DataDog/dd-trace-doe/actions/runs/22102391228/job/63875669025 - During pre-fork, `PeriodicThread.stop()` sets the internal request event to wake the periodic threads promptly. - After fork restart, that same request could be interpreted as a real wakeup, causing an immediate `periodic()` run. - For profiling, that means an immediate `upload()` right after fork, which results in unexpected profile exports. What this PR changfes: - Distinguishes a fork-stop wakeup from a real `awake()` request in `PeriodicThread` - Clears the request after fork only when it came from the pre-fork stop path. - Preserves real `awake()` behavior across restart windows. ## Testing <!-- Describe your testing strategy or note what tests are included --> ## Risks <!-- Note any risks associated with this change, or "None" if no risks --> ## Additional Notes <!-- Any other information that would be helpful for reviewers --> Co-authored-by: taegyun.kim <[email protected]> (cherry picked from commit 295a47f)
…er fork() After #14163, we started to see a profiler behavior change in `dd-trace-doe` CI: - `TestLanguage/python/flush/false` expected `0` profiles, got `3`. https://github.com/DataDog/dd-trace-doe/actions/runs/22102391228/job/63875669025 - During pre-fork, `PeriodicThread.stop()` sets the internal request event to wake the periodic threads promptly. - After fork restart, that same request could be interpreted as a real wakeup, causing an immediate `periodic()` run. - For profiling, that means an immediate `upload()` right after fork, which results in unexpected profile exports. What this PR changfes: - Distinguishes a fork-stop wakeup from a real `awake()` request in `PeriodicThread` - Clears the request after fork only when it came from the pre-fork stop path. - Preserves real `awake()` behavior across restart windows. <!-- Describe your testing strategy or note what tests are included --> <!-- Note any risks associated with this change, or "None" if no risks --> <!-- Any other information that would be helpful for reviewers --> Co-authored-by: taegyun.kim <[email protected]> (cherry picked from commit 295a47f)


We refactor the native periodic thread implementation to be fork-safe whereby all such threads that are running at the time of a fork are automatically stopped before the fork, then restarted after it. We also take care to avoid stopping and restarting threads in the parent process if we detect an immediate call to fork again.
Some of the implications of this change are that there is no longer the need for fork-safe synchronisation objects, such as Locks and Events. This is because all (ddtrace) threads are guaranteed to be stopped before a fork, and restarted afterwards. There is also no longer the need to manually recreate threads, as currently done in many places. The only thing that needs to be taken care of is to ensure that the state of the periodic services is as expected after a fork. For example, many products have the need to avoid sending duplicate values from different processes. This can be achieved by either subclassing from
ForksafeAwakeablePeriodicServiceand implementing theresetmethod to let any periodic service to automatically trigger the reset logic on forks (the preferred way), or by the current approach of registering a fork-safe hook.Performance Analysis
Stopping and restarting threads can be expensive but gives good fork-safety guarantees. Because threads have to be joined before the fork can continue, a forked process might get delayed by a thread currently busy on I/O. Because our periodic threads run on periods that are O(1s), we can expect these unfortunate event to be quite rare. Besides, it is quite likely that a process forks at the very beginning of its execution, where it is unlikely that the periodic threads have had a chance to trigger their first periodic action.
The timer set around close
forkinvocation should reduce delays even more by preventing the parent process from stopping and restarting threads in betweenforkcalls. To give some number, the invocation of 1000 forks in a tight loop with the tracer and the profiler enabled require ~1.2s to complete on an M1. In comparison, the same loop withoutddtracetakes 0.3s. The same loop with stop/restart in between forks takes ~1.5s.The wall-time view through a profiler gives the following picture:
A good portion of the overhead comes from the tracing of
forkitself, which could be improved by moving it to the native layer in follow-up work.Checklist
Reviewer Checklist