Skip to content

New repo layout#3

Merged
masci merged 7 commits into
masterfrom
massi/layout
Jun 10, 2016
Merged

New repo layout#3
masci merged 7 commits into
masterfrom
massi/layout

Conversation

@masci

@masci masci commented Jun 10, 2016

Copy link
Copy Markdown
Contributor

Inspired by kubernetes and etcd, keeps executables and libraries in different packages.
Added shell scripts to build the binaries and run tests.
Massive cleanup of unused resources.

@masci
masci merged commit b6e6a60 into master Jun 10, 2016
@masci
masci deleted the massi/layout branch June 10, 2016 18:22
masci pushed a commit that referenced this pull request Apr 1, 2019
hush-hush added a commit that referenced this pull request Apr 17, 2019
safchain added a commit to safchain/datadog-agent that referenced this pull request May 26, 2020
L3n41c added a commit that referenced this pull request Apr 26, 2021
Fix the following DCA error:
```
2021-04-26 15:23:32 UTC | CLUSTER | ERROR | (pkg/clusteragent/externalmetrics/datadogmetric_controller.go:170 in process) | Impossible to synchronize DatadogMetric (attempt #3): datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: Unable to create DatadogMetric: datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: datadogmetric.datadoghq.com "dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f" is invalid: kind: Invalid value: "datadogmetric": must be DatadogMetric
```
celenechang pushed a commit that referenced this pull request Apr 29, 2021
Fix the following DCA error:
```
2021-04-26 15:23:32 UTC | CLUSTER | ERROR | (pkg/clusteragent/externalmetrics/datadogmetric_controller.go:170 in process) | Impossible to synchronize DatadogMetric (attempt #3): datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: Unable to create DatadogMetric: datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: datadogmetric.datadoghq.com "dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f" is invalid: kind: Invalid value: "datadogmetric": must be DatadogMetric
```
celenechang pushed a commit that referenced this pull request Apr 30, 2021
Fix the following DCA error:
```
2021-04-26 15:23:32 UTC | CLUSTER | ERROR | (pkg/clusteragent/externalmetrics/datadogmetric_controller.go:170 in process) | Impossible to synchronize DatadogMetric (attempt #3): datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: Unable to create DatadogMetric: datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: datadogmetric.datadoghq.com "dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f" is invalid: kind: Invalid value: "datadogmetric": must be DatadogMetric
```
celenechang pushed a commit that referenced this pull request May 5, 2021
Fix the following DCA error:
```
2021-04-26 15:23:32 UTC | CLUSTER | ERROR | (pkg/clusteragent/externalmetrics/datadogmetric_controller.go:170 in process) | Impossible to synchronize DatadogMetric (attempt #3): datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: Unable to create DatadogMetric: datadog-agent-helm/dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f, err: datadogmetric.datadoghq.com "dcaautogen-afda243730e31f83713b52b84d9d2bac9b6d1f" is invalid: kind: Invalid value: "datadogmetric": must be DatadogMetric
```
iglendd added a commit that referenced this pull request Mar 18, 2022
Added reference to the upgraded gohai commit (#3)
einat-stern pushed a commit that referenced this pull request Sep 19, 2022
@akarpz akarpz mentioned this pull request Oct 26, 2023
10 tasks
chouetz pushed a commit that referenced this pull request Apr 23, 2024
* Revert "[acix-24] Notify job status as infra failures using job_failure (#24215)"

This reverts commit 9f09a75.

* Revert "Use gitlab module instead of raw requests (3) (#24867)"

This reverts commit fb56148.
CelianR added a commit that referenced this pull request May 6, 2024
dd-octo-sts Bot added a commit that referenced this pull request Jan 27, 2026
Skip the SSH session patcher and add a test to illustrate the current issue.
In addition, adds the possibility to check specific fields in the json returned for ssh_session events.

### Motivation

The retry mechanism could cause the agent to send no more than one event per minute if an SSH session was not properly resolved.
Previously, the event was not sent and the agent would wait one minute before sending it with the `unknown` type. However, this `authtype` would never be resolved because the session was initialized before the agent started processing events. As a result, every subsequent SSH event would wait one minute for nothing, causing a significant delay in agent events, potentially blocking all the other events.

### Describe how you validated your changes
Added a test that illustrate the issue : `TestSSHUserSessionBlocking`
With this change, the ssh_session event is now sent with `authtype` set to `unknown` and directly sent.

Error without commenting the patcher :
```
        	Error:      	Received unexpected error:
        	            	All attempts fail:
        	            	#1: not found
        	            	#2: not found
        	            	#3: not found
        	            	#4: not found
        	            	#5: not found
        	            	#6: not found
        	            	#7: not found
        	            	#8: not found
        	            	#9: not found
        	            	#10: not found
        	            	#11: not found
        	            	#12: not found
        	            	#13: not found
        	            	#14: not found
        	            	#15: not found
        	            	#16: not found
        	            	#17: not found
        	            	#18: not found
        	            	#19: not found
        	            	#20: not found
        	            	#21: not found
        	            	#22: not found
        	            	#23: not found
        	            	#24: not found
        	            	#25: not found
        	            	#26: not found
        	            	#27: not found
        	            	#28: not found
        	            	#29: not found
        	            	#30: not found
        	Test:       	TestSSHUserSessionBlocking/second_ssh_no_auth
```

Co-authored-by: theo.putegnat <[email protected]>
(cherry picked from commit 40d1f09)

___

Co-authored-by: Théo Putegnat <[email protected]>
gh-worker-dd-mergequeue-cf854d Bot pushed a commit that referenced this pull request Jan 28, 2026
Backport 40d1f09 from #45437.

 ___

### What does this PR do?

Skip the SSH session patcher and add a test to illustrate the current issue.
In addition, adds the possibility to check specific fields in the json returned for ssh_session events.

### Motivation

The retry mechanism could cause the agent to send no more than one event per minute if an SSH session was not properly resolved.
Previously, the event was not sent and the agent would wait one minute before sending it with the `unknown` type. However, this `authtype` would never be resolved because the session was initialized before the agent started processing events. As a result, every subsequent SSH event would wait one minute for nothing, causing a significant delay in agent events, potentially blocking all the other events.

### Describe how you validated your changes
Added a test that illustrate the issue : `TestSSHUserSessionBlocking`
With this change, the ssh_session event is now sent with `authtype` set to `unknown` and directly sent.


Error without commenting the patcher :
```
        	Error:      	Received unexpected error:
        	            	All attempts fail:
        	            	#1: not found
        	            	#2: not found
        	            	#3: not found
        	            	#4: not found
        	            	#5: not found
        	            	#6: not found
        	            	#7: not found
        	            	#8: not found
        	            	#9: not found
        	            	#10: not found
        	            	#11: not found
        	            	#12: not found
        	            	#13: not found
        	            	#14: not found
        	            	#15: not found
        	            	#16: not found
        	            	#17: not found
        	            	#18: not found
        	            	#19: not found
        	            	#20: not found
        	            	#21: not found
        	            	#22: not found
        	            	#23: not found
        	            	#24: not found
        	            	#25: not found
        	            	#26: not found
        	            	#27: not found
        	            	#28: not found
        	            	#29: not found
        	            	#30: not found
        	Test:       	TestSSHUserSessionBlocking/second_ssh_no_auth
```

Co-authored-by: axel.vonengel <[email protected]>
gh-worker-dd-mergequeue-cf854d Bot pushed a commit that referenced this pull request Jan 28, 2026
Backport 40d1f09 from #45437.

 ___

### What does this PR do?

Skip the SSH session patcher and add a test to illustrate the current issue.
In addition, adds the possibility to check specific fields in the json returned for ssh_session events.

### Motivation

The retry mechanism could cause the agent to send no more than one event per minute if an SSH session was not properly resolved.
Previously, the event was not sent and the agent would wait one minute before sending it with the `unknown` type. However, this `authtype` would never be resolved because the session was initialized before the agent started processing events. As a result, every subsequent SSH event would wait one minute for nothing, causing a significant delay in agent events, potentially blocking all the other events.

### Describe how you validated your changes
Added a test that illustrate the issue : `TestSSHUserSessionBlocking`
With this change, the ssh_session event is now sent with `authtype` set to `unknown` and directly sent.


Error without commenting the patcher :
```
        	Error:      	Received unexpected error:
        	            	All attempts fail:
        	            	#1: not found
        	            	#2: not found
        	            	#3: not found
        	            	#4: not found
        	            	#5: not found
        	            	#6: not found
        	            	#7: not found
        	            	#8: not found
        	            	#9: not found
        	            	#10: not found
        	            	#11: not found
        	            	#12: not found
        	            	#13: not found
        	            	#14: not found
        	            	#15: not found
        	            	#16: not found
        	            	#17: not found
        	            	#18: not found
        	            	#19: not found
        	            	#20: not found
        	            	#21: not found
        	            	#22: not found
        	            	#23: not found
        	            	#24: not found
        	            	#25: not found
        	            	#26: not found
        	            	#27: not found
        	            	#28: not found
        	            	#29: not found
        	            	#30: not found
        	Test:       	TestSSHUserSessionBlocking/second_ssh_no_auth
```

Co-authored-by: YoannGh <[email protected]>
Co-authored-by: florent.clarret <[email protected]>
theomagellan pushed a commit that referenced this pull request Feb 2, 2026
### What does this PR do?

Skip the SSH session patcher and add a test to illustrate the current issue.
In addition, adds the possibility to check specific fields in the json returned for ssh_session events.

### Motivation

The retry mechanism could cause the agent to send no more than one event per minute if an SSH session was not properly resolved.
Previously, the event was not sent and the agent would wait one minute before sending it with the `unknown` type. However, this `authtype` would never be resolved because the session was initialized before the agent started processing events. As a result, every subsequent SSH event would wait one minute for nothing, causing a significant delay in agent events, potentially blocking all the other events.

### Describe how you validated your changes
Added a test that illustrate the issue : `TestSSHUserSessionBlocking`
With this change, the ssh_session event is now sent with `authtype` set to `unknown` and directly sent.


Error without commenting the patcher :
```
        	Error:      	Received unexpected error:
        	            	All attempts fail:
        	            	#1: not found
        	            	#2: not found
        	            	#3: not found
        	            	#4: not found
        	            	#5: not found
        	            	#6: not found
        	            	#7: not found
        	            	#8: not found
        	            	#9: not found
        	            	#10: not found
        	            	#11: not found
        	            	#12: not found
        	            	#13: not found
        	            	#14: not found
        	            	#15: not found
        	            	#16: not found
        	            	#17: not found
        	            	#18: not found
        	            	#19: not found
        	            	#20: not found
        	            	#21: not found
        	            	#22: not found
        	            	#23: not found
        	            	#24: not found
        	            	#25: not found
        	            	#26: not found
        	            	#27: not found
        	            	#28: not found
        	            	#29: not found
        	            	#30: not found
        	Test:       	TestSSHUserSessionBlocking/second_ssh_no_auth
```

Co-authored-by: theo.putegnat <[email protected]>
wynbennett added a commit that referenced this pull request Feb 23, 2026
Summary of Changes

  HIGH Priority Issues Fixed:

  #1: Write lock held across network I/O (impl/delegatedauth.go:270)
  - Refactored refreshAndGetAPIKey to release the lock before making network calls (authenticate)
  - The lock is now only held briefly to check/update state, not during network I/O

  #2: Context not propagated to signer.SignHTTP (aws.go:195)
  - Updated generateAwsAuthData to accept a context parameter
  - Changed signer.SignHTTP(context.Background(), ...) to signer.SignHTTP(ctx, ...)

  #3: Context not propagated to getCredentials IMDS call (aws.go:119)
  - Updated getCredentials to accept a context parameter
  - Removed ctx := context.Background() and now uses the passed context for IMDS calls

  MEDIUM Priority Issues Fixed:

  #4: No response body size limit (api/delegated_auth.go:97)
  - Added maxResponseBodySize = 1 * 1024 * 1024 constant (1 MB)
  - Wrapped response body with io.LimitReader to prevent memory exhaustion

  #5: No overall HTTP client timeout (api/delegated_auth.go:82)
  - Added httpClientTimeout = 30 * time.Second constant
  - Added Timeout: httpClientTimeout to the HTTP client

  #6: config.Set called while holding write lock (impl/delegatedauth.go:341)
  - Moved updateConfigWithAPIKey call outside the lock in startBackgroundRefresh
  - Captured the API key while holding the lock, then released it before calling config.Set

  #7: Blocking IMDS calls while holding write lock (impl/delegatedauth.go:127)
  - Refactored initializeIfNeeded to perform cloud detection without holding locks
  - IMDS calls now happen outside any lock, then state is updated with a brief write lock

  #8: Regex fails silently for non-standard formats (api/delegated_auth.go:36)
  - Added debug log when endpoint doesn't match known Datadog domain pattern
  - Updated function documentation to clarify behavior

  #9: Uncached IMDS credential fetch (aws.go:104)
  - Added documentation explaining the trade-off (refresh interval is typically 60 minutes, so caching is not critical)

  #10: Auth proof format undocumented (aws.go:98)
  - Added detailed comment documenting the auth proof format: <base64-body>|<base64-headers>|<method>|<base64-url>

  LOW Priority Issues Fixed:

  #11: Unnecessarily exported types (aws.go)
  - Changed SigningData to signingData (unexported)
  - Changed AWSAuth.AwsRegion to AWSAuth.region (unexported)
  - Updated all references in aws.go and aws_test.go

  #12: Tests exercise copy of goroutine (impl/delegatedauth_test.go:19)
  - Added documentation explaining why tests use a simplified goroutine pattern
  - Clarified that integration tests cover the actual startBackgroundRefresh function

  #13: Subsequent Config param silently ignored (def/delegatedauth.go:24)
  - Updated documentation to clearly state that only the first Config is used
  - Added warning log when a different Config is passed on subsequent calls
JSGette added a commit that referenced this pull request Apr 15, 2026
misteriaud added a commit that referenced this pull request Apr 16, 2026
…ctions

Replace the monolithic batcher (5 ring buffers sharing one transport)
with a generic pipeline[T] struct. Each pipeline owns its own ring
buffers, flush goroutines, and dedicated UDS connection:

  metricsPipeline = pipeline[metricPoint]        + unixConn #1
  logsPipeline    = pipeline[logEntry]            + unixConn #2
  tracePipeline   = pipeline[capturedTraceStat]   + unixConn #3

Pipelines are fully independent — one slow pipeline (e.g. logs sending
large frames) cannot block or starve another.

Key changes:
- pipeline[T]: generic struct with AddEntry(T), AddContextDef, Stop.
  1-2 flush goroutines per pipeline (entries + optional contexts).
  flushChunked reused unchanged.
- unixConn: simple per-connection transport replacing pooledTransport.
  Lazy dial, mutex held during Send, reconnect once on error.
- activate(): creates 3 pipelines with 3 independent connections.
  sync.Once coordinates teardown when any transport disconnects.

Testbench results (all 0 drops):
- dogstatsd-p99: 2.8M metrics sent
- logs-high-throughput: 40M logs at 10 MiB/s, 3 GB Parquet
- metrics-logs-combined: 743K metrics + 1M logs

Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
victorsprengel added a commit that referenced this pull request Jun 3, 2026
Extends three test files:
- packet/packet_test.go: TestGetTagsWithCustomTags{,SNMPV1} verify the
  built-in tags stay first and user tags append in declared order;
  TestGetTagsWithEmptyCustomTags covers the empty/nil path (PRD success
  criterion #4 — byte-for-byte unchanged behaviour).
- config/def/config_test.go: covers YAML unmarshal, default-empty,
  whitespace/empty normalization, and the OversizedTags helper.
- formatter/impl/formatter_test.go: TestFormatPacketIncludesCustomTagsInDDTags
  asserts ddtags includes user tags (criterion #2);
  TestFormatPacketCustomTagsOnTelemetry asserts traps_not_enriched and
  incorrect_format both carry the custom tags (criterion #3). Split out
  TestFormatPacketToJSONBody to preserve the original variable-by-variable
  coverage now that TestFormatPacketToJSON only covers ddtags.

Co-Authored-By: Claude <[email protected]>
celenechang added a commit that referenced this pull request Jun 10, 2026
Replaces the brittle 'HorizontalLastActions > 0' gate with an explicit
LastScaledTarget tracker on PodAutoscalerInternal. The tracker records
(namespace, name, GVK) every time the horizontal controller successfully
writes `.spec.replicas`, and is cleared after a successful release. The
release path now fires on three triggers, not one:

  - horizontal scaling disabled  (existing)
  - apply mode switched to Preview  (#4 — previously not covered)
  - the DPA's TargetRef was retargeted to a different workload, so the
    OLD target still holds the stale managedFields entry while the new
    target has none  (#7 — previously released against the wrong target)

Release-failure handling also overhauled:

  - On failure, the helper now constructs a ConditionError, emits a
    Warning Event on the DPA, calls UpdateFromHorizontalAction(nil, err),
    increments HorizontalActionErrorInc, and returns the error so the
    workqueue's maxRetry guard caps the loop instead of hot-looping
    invisibly with (Requeue, nil)  (#3 — was previously silent).
  - The three delete branches (remote-owned, profile-managed,
    local-owned) now retry release up to maxRetry attempts via
    c.Workqueue.NumRequeues; on exhaustion they log Errorf and proceed
    with the delete so a permanently broken release (RBAC never granted)
    cannot indefinitely block a user from deleting a DPA  (#1).

Tests:
  - Existing TestHorizontalControllerReleaseOwnershipOnDisable updated
    to seed LastScaledTarget instead of HorizontalLastActions.
  - New TestHorizontalControllerReleaseOwnershipOnPreviewTransition.
  - New TestHorizontalControllerReleaseOwnershipOnTargetRefChange
    asserts release fires against the OLD target, not the spec's
    current target.
  - TestHorizontalControllerReleaseOwnershipOnDisable_FailureRetainsState
    rewritten to assert (Requeue, err) and LastScaledTarget retention.
  - TestLeaderCreateDeleteLocal / TestLeaderCreateDeleteRemote updated
    to seed LastScaledTarget so the delete path actually exercises the
    release call.
  - testScalingDecision now mirrors SetLastScaledTarget after every
    successful scale, matching production semantics.

LastScaledTarget is in-memory only — on cluster-agent restart it resets
and the next successful scale re-populates it. The trade-off is a
narrow window where a DPA disabled across a controller restart with no
subsequent scale would not release; acceptable in exchange for not
needing CRD status schema changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
celenechang added a commit that referenced this pull request Jun 10, 2026
Replaces the brittle 'HorizontalLastActions > 0' gate with an explicit
LastScaledTarget tracker on PodAutoscalerInternal. The tracker records
(namespace, name, GVK) every time the horizontal controller successfully
writes `.spec.replicas`, and is cleared after a successful release. The
release path now fires on three triggers, not one:

  - horizontal scaling disabled  (existing)
  - apply mode switched to Preview  (#4 — previously not covered)
  - the DPA's TargetRef was retargeted to a different workload, so the
    OLD target still holds the stale managedFields entry while the new
    target has none  (#7 — previously released against the wrong target)

Release-failure handling also overhauled:

  - On failure, the helper now constructs a ConditionError, emits a
    Warning Event on the DPA, calls UpdateFromHorizontalAction(nil, err),
    increments HorizontalActionErrorInc, and returns the error so the
    workqueue's maxRetry guard caps the loop instead of hot-looping
    invisibly with (Requeue, nil)  (#3 — was previously silent).
  - The three delete branches (remote-owned, profile-managed,
    local-owned) now retry release up to maxRetry attempts via
    c.Workqueue.NumRequeues; on exhaustion they log Errorf and proceed
    with the delete so a permanently broken release (RBAC never granted)
    cannot indefinitely block a user from deleting a DPA  (#1).

Tests:
  - Existing TestHorizontalControllerReleaseOwnershipOnDisable updated
    to seed LastScaledTarget instead of HorizontalLastActions.
  - New TestHorizontalControllerReleaseOwnershipOnPreviewTransition.
  - New TestHorizontalControllerReleaseOwnershipOnTargetRefChange
    asserts release fires against the OLD target, not the spec's
    current target.
  - TestHorizontalControllerReleaseOwnershipOnDisable_FailureRetainsState
    rewritten to assert (Requeue, err) and LastScaledTarget retention.
  - TestLeaderCreateDeleteLocal / TestLeaderCreateDeleteRemote updated
    to seed LastScaledTarget so the delete path actually exercises the
    release call.
  - testScalingDecision now mirrors SetLastScaledTarget after every
    successful scale, matching production semantics.

LastScaledTarget is in-memory only — on cluster-agent restart it resets
and the next successful scale re-populates it. The trade-off is a
narrow window where a DPA disabled across a controller restart with no
subsequent scale would not release; acceptable in exchange for not
needing CRD status schema changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
celenechang added a commit that referenced this pull request Jun 10, 2026
Replaces the brittle 'HorizontalLastActions > 0' gate with an explicit
LastScaledTarget tracker on PodAutoscalerInternal. The tracker records
(namespace, name, GVK) every time the horizontal controller successfully
writes `.spec.replicas`, and is cleared after a successful release. The
release path now fires on three triggers, not one:

  - horizontal scaling disabled  (existing)
  - apply mode switched to Preview  (#4 — previously not covered)
  - the DPA's TargetRef was retargeted to a different workload, so the
    OLD target still holds the stale managedFields entry while the new
    target has none  (#7 — previously released against the wrong target)

Release-failure handling also overhauled:

  - On failure, the helper now constructs a ConditionError, emits a
    Warning Event on the DPA, calls UpdateFromHorizontalAction(nil, err),
    increments HorizontalActionErrorInc, and returns the error so the
    workqueue's maxRetry guard caps the loop instead of hot-looping
    invisibly with (Requeue, nil)  (#3 — was previously silent).
  - The three delete branches (remote-owned, profile-managed,
    local-owned) now retry release up to maxRetry attempts via
    c.Workqueue.NumRequeues; on exhaustion they log Errorf and proceed
    with the delete so a permanently broken release (RBAC never granted)
    cannot indefinitely block a user from deleting a DPA  (#1).

Tests:
  - Existing TestHorizontalControllerReleaseOwnershipOnDisable updated
    to seed LastScaledTarget instead of HorizontalLastActions.
  - New TestHorizontalControllerReleaseOwnershipOnPreviewTransition.
  - New TestHorizontalControllerReleaseOwnershipOnTargetRefChange
    asserts release fires against the OLD target, not the spec's
    current target.
  - TestHorizontalControllerReleaseOwnershipOnDisable_FailureRetainsState
    rewritten to assert (Requeue, err) and LastScaledTarget retention.
  - TestLeaderCreateDeleteLocal / TestLeaderCreateDeleteRemote updated
    to seed LastScaledTarget so the delete path actually exercises the
    release call.
  - testScalingDecision now mirrors SetLastScaledTarget after every
    successful scale, matching production semantics.

LastScaledTarget is in-memory only — on cluster-agent restart it resets
and the next successful scale re-populates it. The trade-off is a
narrow window where a DPA disabled across a controller restart with no
subsequent scale would not release; acceptable in exchange for not
needing CRD status schema changes.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
JSGette added a commit that referenced this pull request Jun 11, 2026
BarFinsdd added a commit that referenced this pull request Jul 22, 2026
Two changes driven by comparing our reconstruction to how the dd-trace-go SDK
presents LLM Obs.

Conversation-thread grouping (replaces the earlier session-agent wrapper):
- A long-lived agent span now groups every turn of one *conversation*, keyed by
  threadKey = session_id + the conversation's first user message (re-sent on
  every follow-up turn). All the LLM/tool calls of a multi-turn conversation
  nest under one agent flow; session_id stays as a tag so the UI's Sessions view
  still groups conversations. This matches the SDK model (an agent = one
  execution; session_id groups traces) rather than conflating a whole session
  into one agent. The agent is closed by an idle reaper (llmConvTTL) since the
  wire has no conversation-end signal.
- Known wire limitation: two conversations with an identical first prompt in the
  same session collapse to one thread (documented on threadKey).

Embedding spans (#3):
- A captured body with "input" and no "messages" is an /v1/embeddings call; it
  now emits an embedding-kind span (input document + token usage) instead of a
  chat llm span. Detection is on the body since the consumer only has the DATA
  frame.

Tests: TestThreadKeyGroupsTurns (turns of one conversation share a key; distinct
prompts/sessions don't) and TestIsEmbeddingBody. All existing pairing,
reassembly, and parser tests still pass. Temporary diagnostic logs removed.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant