fix(onboard-skills): keep emoji / surrogate pairs intact during skill description truncation#101619
fix(onboard-skills): keep emoji / surrogate pairs intact during skill description truncation#101619maweibin wants to merge 1 commit into
Conversation
… description truncation Use truncateUtf16Safe() instead of .slice() to count full code points, preventing lone surrogates in onboarding skill summary output. Co-Authored-By: Claude Opus 4.8 <[email protected]>
|
Codex review: needs real behavior proof before merge. Reviewed July 7, 2026, 8:14 AM ET / 12:14 UTC. Summary PR surface: Source +1, Tests +9. Total +10 across 2 files. Reproducibility: yes. at source level: current main uses Review metrics: none identified. Merge readiness Overall follows the weaker of proof and patch quality, so missing proof can cap an otherwise strong patch. Rank-up moves:
Proof guidance:
Risk before merge
Maintainer options:
Next step before merge
Security Review findings
Review detailsBest possible solution: Land the existing-helper production fix after the PR adds redacted real onboarding output proof and a regression test that drives Do we have a high-confidence way to reproduce the issue? Yes at source level: current main uses Is this the best way to solve the issue? Yes for the production change: using the existing Full review comments:
Overall correctness: patch is correct AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 2ba622ca3019. Label changesLabel justifications:
Evidence reviewedPR surface: Source +1, Tests +9. Total +10 across 2 files. View PR surface stats
What I checked:
Likely related people:
What the crustacean ranks mean
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics. How this review workflow works
Review history (1 earlier review cycle)
|
|
Thanks @maweibin. The useful behavior from this PR was consolidated with the sibling UTF-16 boundary fixes into #101654 and landed on The landed fix uses the existing |
What Problem This Solves
summarizeInstallFailure / formatSkillHintinsrc/commands/onboard-skills.tstruncates text with.slice()which counts UTF-16 code units. When a surrogate-pair character (emoji) straddles the cut boundary,.slice()produces a lone surrogate (e.g.\ud83d) that renders as `` in terminal output.This is user-visible: the truncated text reaches user-facing CLI, status, or log output, so any content containing emoji near the 139-character boundary can hit this.
Why This Change Was Made
Replace
.slice()withtruncateUtf16Safe()from@openclaw/normalization-core/utf16-sliceso the truncation counts full code points instead of UTF-16 code units. This matches the already-merged PR #101517 (session-cost-usage.ts) and the canonicalshortenTexthelper insrc/commands/text-format.ts:5.User Impact
Users will no longer see `` replacement characters in this output when emoji appear near truncation boundaries. Existing column width / character caps are preserved — the broken surrogate is dropped cleanly rather than rendered as a replacement character.
Evidence
Real environment tested: local OpenClaw source checkout, Node v24.13.1, PR head.
Exact steps after this patch:
Evidence after fix:
Observed result after fix:
truncateUtf16Safepreserves full code points; no lone surrogate reaches output. The broken surrogate pair is dropped cleanly rather than producing ``.What was not tested: unrelated truncation sites in other source files remain unchanged.
Regression Test Plan
Includes a focused regression test that verifies surrogate pair preservation at the truncation boundary.
AI-assisted.