[Serverless-Init] Cloud Run Jobs Inferred Span#43371
Conversation
Static quality checks✅ Please find below the results from static quality gates Successful checksInfo
|
Regression DetectorRegression Detector ResultsMetrics dashboard Baseline: 334d0f1 Optimization Goals: ✅ No significant changes detected
|
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | docker_containers_cpu | % cpu utilization | +1.45 | [-1.48, +4.39] | 1 | Logs |
Fine details of change detection per experiment
| perf | experiment | goal | Δ mean % | Δ mean % CI | trials | links |
|---|---|---|---|---|---|---|
| ➖ | quality_gate_logs | % cpu utilization | +4.62 | [+3.10, +6.13] | 1 | Logs bounds checks dashboard |
| ➖ | docker_containers_cpu | % cpu utilization | +1.45 | [-1.48, +4.39] | 1 | Logs |
| ➖ | docker_containers_memory | memory utilization | +0.71 | [+0.50, +0.93] | 1 | Logs |
| ➖ | quality_gate_metrics_logs | memory utilization | +0.57 | [+0.37, +0.77] | 1 | Logs bounds checks dashboard |
| ➖ | uds_dogstatsd_20mb_12k_contexts_20_senders | memory utilization | +0.48 | [+0.43, +0.53] | 1 | Logs |
| ➖ | ddot_metrics_sum_delta | memory utilization | +0.40 | [+0.19, +0.61] | 1 | Logs |
| ➖ | ddot_metrics_sum_cumulativetodelta_exporter | memory utilization | +0.19 | [-0.05, +0.42] | 1 | Logs |
| ➖ | file_tree | memory utilization | +0.15 | [+0.10, +0.20] | 1 | Logs |
| ➖ | file_to_blackhole_500ms_latency | egress throughput | +0.07 | [-0.31, +0.45] | 1 | Logs |
| ➖ | file_to_blackhole_0ms_latency | egress throughput | +0.04 | [-0.38, +0.45] | 1 | Logs |
| ➖ | otlp_ingest_logs | memory utilization | +0.02 | [-0.07, +0.12] | 1 | Logs |
| ➖ | file_to_blackhole_100ms_latency | egress throughput | +0.01 | [-0.04, +0.06] | 1 | Logs |
| ➖ | uds_dogstatsd_to_api | ingress throughput | +0.00 | [-0.13, +0.13] | 1 | Logs |
| ➖ | tcp_dd_logs_filter_exclude | ingress throughput | +0.00 | [-0.07, +0.08] | 1 | Logs |
| ➖ | file_to_blackhole_1000ms_latency | egress throughput | -0.00 | [-0.42, +0.41] | 1 | Logs |
| ➖ | uds_dogstatsd_to_api_v3 | ingress throughput | -0.00 | [-0.13, +0.13] | 1 | Logs |
| ➖ | otlp_ingest_metrics | memory utilization | -0.04 | [-0.19, +0.10] | 1 | Logs |
| ➖ | quality_gate_idle_all_features | memory utilization | -0.06 | [-0.11, -0.01] | 1 | Logs bounds checks dashboard |
| ➖ | quality_gate_idle | memory utilization | -0.07 | [-0.11, -0.02] | 1 | Logs bounds checks dashboard |
| ➖ | ddot_metrics | memory utilization | -0.46 | [-0.68, -0.25] | 1 | Logs |
| ➖ | ddot_metrics_sum_cumulative | memory utilization | -0.62 | [-0.77, -0.48] | 1 | Logs |
| ➖ | ddot_logs | memory utilization | -0.72 | [-0.79, -0.66] | 1 | Logs |
| ➖ | tcp_syslog_to_blackhole | ingress throughput | -1.46 | [-1.53, -1.39] | 1 | Logs |
Bounds Checks: ✅ Passed
| perf | experiment | bounds_check_name | replicates_passed | links |
|---|---|---|---|---|
| ✅ | docker_containers_cpu | simple_check_run | 10/10 | |
| ✅ | docker_containers_memory | memory_usage | 10/10 | |
| ✅ | docker_containers_memory | simple_check_run | 10/10 | |
| ✅ | file_to_blackhole_0ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_0ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_1000ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_1000ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_100ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_100ms_latency | memory_usage | 10/10 | |
| ✅ | file_to_blackhole_500ms_latency | lost_bytes | 10/10 | |
| ✅ | file_to_blackhole_500ms_latency | memory_usage | 10/10 | |
| ✅ | quality_gate_idle | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_idle_all_features | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | lost_bytes | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_logs | memory_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | cpu_usage | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | intake_connections | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | lost_bytes | 10/10 | bounds checks dashboard |
| ✅ | quality_gate_metrics_logs | memory_usage | 10/10 | bounds checks dashboard |
Explanation
Confidence level: 90.00%
Effect size tolerance: |Δ mean %| ≥ 5.00%
Performance changes are noted in the perf column of each table:
- ✅ = significantly better comparison variant performance
- ❌ = significantly worse comparison variant performance
- ➖ = no significant change in performance
A regression test is an A/B test of target performance in a repeatable rig, where "performance" is measured as "comparison variant minus baseline variant" for an optimization goal (e.g., ingress throughput). Due to intrinsic variability in measuring that goal, we can only estimate its mean value for each experiment; we report uncertainty in that value as a 90.00% confidence interval denoted "Δ mean % CI".
For each experiment, we decide whether a change in performance is a "regression" -- a change worth investigating further -- if all of the following criteria are true:
-
Its estimated |Δ mean %| ≥ 5.00%, indicating the change is big enough to merit a closer look.
-
Its 90.00% confidence interval "Δ mean % CI" does not contain zero, indicating that if our statistical model is accurate, there is at least a 90.00% chance there is a difference in performance between baseline and comparison variants.
-
Its configuration does not mark it "erratic".
CI Pass/Fail Decision
✅ Passed. All Quality Gates passed.
- quality_gate_logs, bounds check lost_bytes: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_idle_all_features, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check cpu_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check intake_connections: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check memory_usage: 10/10 replicas passed. Gate passed.
- quality_gate_metrics_logs, bounds check lost_bytes: 10/10 replicas passed. Gate passed.
|
|
||
| // initJobSpan creates and initializes the job span with Cloud Run Job metadata | ||
| func (c *CloudRunJobs) initJobSpan() { | ||
| tags := c.GetTags() |
There was a problem hiding this comment.
Since we have changes to tags planned that will result in a different set of tags for the traceService vs the logsService and metricsService, might it make sense to get span tags from the trace service? That will include configured tags and base tags as well; I don't know if you were purposefully avoiding those. It looks like the traceService has a fair bit of separate tag logic.
There was a problem hiding this comment.
Yes, once we separate the tags, that would simplify this code a bit. I'm happy to wait for that to merge and resolve conflicts in this PR, or merge this PR first and help resolve conflicts in that PR
kathiehuang
left a comment
There was a problem hiding this comment.
LGTM! Just curious, did you reference the Lambda extension or anything when building out this logic?
Yes, for span creation logic, I referenced the Go version of the Lambda Extension. I looked at the commit before this PR: #42994 and I simplified/adapted the code a bit, since most of that code was several years old |
|
/merge |
|
View all feedbacks in Devflow UI.
The expected merge time in
|
What does this PR do?
Create an inferred span for the duration of the task execution.

View span
And it correctly reparents spans received from the tracer (custom, autoinstrumentation, etc.)

This can be disabled with the env var
DD_APM_ENABLED=falseMotivation
So traces work OOTB, no custom span creation is required from the user. This also follows the convention set by Lambda.
As an added benefit, we can get rid of the high cardinality tags on the existing enhanced metrics. Instead, we can just rely on this trace existing to display the duration of each task for a specific execution ID on the frontend. (This will be done in a future PR, and it requires frontend changes)
Describe how you validated your changes
Testing manually with a real Cloud Run Job
Additional Notes
While it's not ideal to have APM logic in the agent, there is no other way to do this. In Lambda, the tracers can instrument the Lambda handler function. But in Cloud Run Jobs, there is no such entry point to instrument. The tracer only knows it's in a container, but not that it's in a Cloud Run Job.
We have a precedent for creating spans in the agent (in the Lambda Extension), and there were no issues there.