fix(audit): keep emoji / surrogate pairs intact during audit text truncation#101628
fix(audit): keep emoji / surrogate pairs intact during audit text truncation#101628maweibin wants to merge 1 commit into
Conversation
…ncation Use truncateUtf16Safe() instead of .slice() to count full code points, preventing lone surrogates in audit CLI output. Co-Authored-By: Claude Opus 4.8 <[email protected]>
|
Codex review: needs real behavior proof before merge. Reviewed July 7, 2026, 8:15 AM ET / 12:15 UTC. Summary PR surface: Source +3, Tests +13. Total +16 across 2 files. Reproducibility: yes. by source inspection: current main sanitizes audit text and then truncates with UTF-16 Review metrics: none identified. Merge readiness Overall follows the weaker of proof and patch quality, so missing proof can cap an otherwise strong patch. Rank-up moves:
Proof guidance:
Risk before merge
Maintainer options:
Next step before merge
Security Review findings
Review detailsBest possible solution: Land the narrow audit truncation fix after the regression proof exercises Do we have a high-confidence way to reproduce the issue? Yes, by source inspection: current main sanitizes audit text and then truncates with UTF-16 Is this the best way to solve the issue? Yes for the implementation shape: using the existing Full review comments:
Overall correctness: patch is correct AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 2ba622ca3019. Label changesLabel changes:
Label justifications:
Evidence reviewedPR surface: Source +3, Tests +13. Total +16 across 2 files. View PR surface stats
What I checked:
Likely related people:
What the crustacean ranks mean
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics. How this review workflow works
|
|
Thanks @maweibin. The useful behavior from this PR was consolidated with the sibling UTF-16 boundary fixes into #101654 and landed on The landed fix uses the existing |
What Problem This Solves
the audit text shortenerinsrc/commands/audit.tstruncates text with.slice()which counts UTF-16 code units. When a surrogate-pair character (emoji) straddles the cut boundary,.slice()produces a lone surrogate (e.g.\ud83d) that renders as�in terminal output.This is user-visible: the truncated text comes from real user-facing output, so any content containing emoji near the 78-character boundary can hit this.
Why This Change Was Made
Replace
.slice()withtruncateUtf16Safe()from@openclaw/normalization-core/utf16-sliceso the truncation counts full code points instead of UTF-16 code units. This matches the already-merged PR #101517 (session-cost-usage) and the existingshortenTexthelper insrc/commands/text-format.ts:5.User Impact
Users will no longer see
�replacement characters in this output when emoji appear near truncation boundaries.Evidence
Real environment tested: local OpenClaw source checkout, Node v24.13.1, PR head.
Exact steps after this patch:
Evidence after fix:
Observed result after fix:
truncateUtf16Safepreserves full code points; no lone surrogate reaches output.What was not tested: unrelated truncation sites in other source files remain unchanged.
Regression Test Plan
AI-assisted.