Separation of tags between metrics logs and traces fix#43370
Conversation
Static quality checks✅ Please find below the results from static quality gates Successful checksInfo
|
Regression DetectorRegression Detector ResultsMetrics dashboard Baseline: bb99c82 Optimization Goals: ✅ No significant changes detected
|
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | docker_containers_cpu | % cpu utilization | +0.47 | [-2.57, +3.52] | 1 | Logs |
Fine details of change detection per experiment
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | quality_gate_metrics_logs | memory utilization | +1.63 | [+1.41, +1.86] | 1 | Logs bounds checks dashboard |
| ➖ | ddot_logs | memory utilization | +0.81 | [+0.73, +0.88] | 1 | Logs |
| ➖ | quality_gate_idle_all_features | memory utilization | +0.53 | [+0.50, +0.57] | 1 | Logs bounds checks dashboard |
| ➖ | docker_containers_cpu | % cpu utilization | +0.47 | [-2.57, +3.52] | 1 | Logs |
| ➖ | uds_dogstatsd_20mb_12k_contexts_20_senders | memory utilization | +0.33 | [+0.27, +0.38] | 1 | Logs |
| ➖ | ddot_metrics_sum_cumulativetodelta_exporter | memory utilization | +0.06 | [-0.17, +0.29] | 1 | Logs |
| ➖ | file_to_blackhole_1000ms_latency | egress throughput | +0.04 | [-0.39, +0.46] | 1 | Logs |
| ➖ | uds_dogstatsd_to_api_v3 | ingress throughput | +0.01 | [-0.11, +0.13] | 1 | Logs |
| ➖ | file_to_blackhole_0ms_latency | egress throughput | +0.00 | [-0.41, +0.41] | 1 | Logs |
| ➖ | tcp_dd_logs_filter_exclude | ingress throughput | -0.00 | [-0.08, +0.07] | 1 | Logs |
| ➖ | uds_dogstatsd_to_api | ingress throughput | -0.00 | [-0.13, +0.12] | 1 | Logs |
| ➖ | file_to_blackhole_100ms_latency | egress throughput | -0.02 | [-0.07, +0.03] | 1 | Logs |
| ➖ | file_to_blackhole_500ms_latency | egress throughput | -0.05 | [-0.43, +0.32] | 1 | Logs |
| ➖ | ddot_metrics_sum_delta | memory utilization | -0.06 | [-0.26, +0.15] | 1 | Logs |
| ➖ | quality_gate_idle | memory utilization | -0.10 | [-0.14, -0.05] | 1 | Logs bounds checks dashboard |
| ➖ | ddot_metrics | memory utilization | -0.12 | [-0.35, +0.11] | 1 | Logs |
| ➖ | file_tree | memory utilization | -0.24 | [-0.31, -0.17] | 1 | Logs |
| ➖ | otlp_ingest_logs | memory utilization | -0.27 | [-0.37, -0.17] | 1 | Logs |
| ➖ | otlp_ingest_metrics | memory utilization | -0.32 | [-0.47, -0.16] | 1 | Logs |
| ➖ | docker_containers_memory | memory utilization | -0.42 | [-0.50, -0.35] | 1 | Logs |
| ➖ | quality_gate_logs | % cpu utilization | -0.46 | [-1.94, +1.02] | 1 | Logs bounds checks dashboard |
| ➖ | tcp_syslog_to_blackhole | ingress throughput | -0.69 | [-0.77, -0.61] | 1 | Logs |
| ➖ | ddot_metrics_sum_cumulative | memory utilization | -0.85 | [-1.01, -0.69] | 1 | Logs |
Bounds Checks: ✅ Passed
| perf | experiment | bounds_check_name | replicates_passed | links |
|---|---|---|---|---|
| ✅ | docker_containers_cpu | simple_check_run | 10/10 | |
| ✅ | docker_containers_memory | memory_usage | 10/10 | |
| ✅ | docker_containers_memory | simple_check_run | 10/10 | |
| ✅ | file_to_blackhole_0ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_0ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_1000ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_1000ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_100ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_100ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_500ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_500ms_latency | memory_usage | 10/10 | |
| ✅ | quality_gate_idle | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | lost_bytes | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | cpu_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | lost_bytes | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | memory_usage | 10/10 | bounds checks dashboard |
Explanation
Confidence level: 90.00%
Effect size tolerance: |Δ mean %| ≥ 5.00%
Performance changes are noted in the perf column of each table:
- ✅ = significantly better comparison variant performance
- ❌ = significantly worse comparison variant performance
- ➖ = no significant change in performance
A regression test is an A/B test of target performance in a repeatable rig, where "performance" is measured as "comparison variant minus baseline variant" for an optimization goal (e.g., ingress throughput). Due to intrinsic variability in measuring that goal, we can only estimate its mean value for each experiment; we report uncertainty in that value as a 90.00% confidence interval denoted "Δ mean % CI".
For each experiment, we decide whether a change in performance is a "regression" -- a change worth investigating further -- if all of the following criteria are true:
-
Its estimated |Δ mean %| ≥ 5.00%, indicating the change is big enough to merit a closer look.
-
Its 90.00% confidence interval "Δ mean % CI" does not contain zero, indicating that if our statistical model is accurate, there is at least a 90.00% chance there is a difference in performance between baseline and comparison variants.
-
Its configuration does not mark it "erratic".
CI Pass/Fail Decision
✅ Passed. All Quality Gates passed.
- quality_gate_idle, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check lost_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check lost_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
54ab478 to
3a58bff
Compare
…43837) ### What does this PR do? This is a temporary fix to remove high cardinality tags from Cloud Run Jobs enhanced metrics. A more thorough refactor will be done in a future PR (#43370) High cardinality removed: <img width="400" height="393" alt="Screenshot 2025-12-04 at 2 11 05 PM" src="https://github.com/user-attachments/assets/0d002d7f-5d58-448d-ae24-8a9c62bf5849" /> Low cardinality still exists: <img width="400" height="382" alt="Screenshot 2025-12-04 at 2 11 39 PM" src="https://github.com/user-attachments/assets/076e73c6-769b-4f11-b023-b2579905a9c7" /> ### Motivation Cost savings ### Describe how you validated your changes Manually with a real Cloud Run Job ### Additional Notes No changelog because Cloud Run Jobs support is not officially released yet. Co-authored-by: nicholas.hulston <[email protected]>
|
Hi there, thanks for this PR. I've added a |
1c262e0 to
ae56b67
Compare
apiarian-datadog
left a comment
There was a problem hiding this comment.
how are we testing this? should we have some kind of self-monitoring alert for the resulting tagged things?
ca24a7d to
c7cf356
Compare
@apiarian-datadog There's a checklist for manual tests in the PR description under |
| @@ -123,10 +123,11 @@ func setup(secretComp secrets.Component, _ mode.Conf, tagger tagger.Component, c | |||
|
|
|||
| log.Debugf("Detected cloud service: %s", cloudService.GetOrigin()) | |||
|
|
|||
| configuredTags := configUtils.GetConfiguredTags(pkgconfigsetup.Datadog(), false) | |||
| tags := serverlessInitTag.GetBaseTagsMapWithMetadata( | |||
There was a problem hiding this comment.
nit: while we're in here, might be worth renaming "tags" to "base tags" or something like that, so it's clear that we probably shouldn't use it directly without serious thought.
i think the existing trace stats tests should cover your trace stats needs on this one. i suppose it might be worth adding a small test that creates an app with some tags and confirms that those tags (and the additional automatic ones) end up on the metrics, traces, and logs that we ingest. or perhaps this is an element of self-monitoring. |
0cc5213 to
13a353a
Compare
Go Package Import DifferencesBaseline: 2441ad4
|
d0e1091 to
ff7d6c5
Compare
ff7d6c5 to
38fe7b9
Compare
3af4456 to
702b72e
Compare
…hMetadata to isolate _dd.compute_stats to metrics
…r to match tag simplification in serverless-init
702b72e to
080a59d
Compare
|
/merge |
|
View all feedbacks in Devflow UI.
The expected merge time in
|
What does this PR do?
_dd.compute_statsis a marker for agent-side trace stats computation. We should isolate_dd.compute_statsto theserverless-inittraces agent.DD_SERVERLESS_INIT_ENABLE_BACKEND_TRACE_STATSandDD_SERVERLESS_INIT_DISABLE_TRACE_STATSMotivation
_dd.compute_statsshould not show on logs or as a tag available for runtime metrics.DD_SERVERLESS_INIT_ENABLE_BACKEND_TRACE_STATSis true,_dd.compute_statsshould show on traces.DD_SERVERLESS_INIT_ENABLE_BACKEND_TRACE_STATSis true andDD_SERVERLESS_INIT_DISABLE_TRACE_STATSis false, we should have accurate metric/trace data.DD_SERVERLESS_INIT_ENABLE_BACKEND_TRACE_STATSis true andDD_SERVERLESS_INIT_DISABLE_TRACE_STATSis false, AND sampling is <1.0, we should have inaccurate metric trace data.Describe how you validated your changes
cloudruncloudrunjobsappservicecontainerappUsing the
lewis/SVLS-4573/serverless-init-test-compute-statsself-monitoring branchcloudrunin self-monitoringappservicein self-monitoringcontainerappin self-monitoringTesting changes to the Azure App Services Extension