Repository navigation
Fix #2469: [Bug] memos-local-plugin: capture summarizer writes Chinese summaries for Englis - #2471
Open
Memtensor-AI wants to merge 2 commits into
Open
Memtensor-AI wants to merge 2 commits into
Memtensor-AI wants to merge 2 commits into
Conversation
…Tensor#2469) The capture SYSTEM_PROMPT anchored its language rule on an ambiguous referent ("the user's original language") and carried a lone CJK example ("用户说了") that primed some models (observed with openai/gpt-6-luna at temp 0) to emit Chinese summaries for English conversations — up to 89% CJK on tool-call summaries and 66% on conversational summaries. Fix per issue MemTensor#2469 (reporter-validated on the same model, 0/30 CJK on English inputs after the change): - Replace "in the user's original language" with "written in the same language as the USER text (English text gets an English summary)". - Remove the Chinese example from the "Do NOT prefix" rule, keeping only "The user said". Also adds a regression unit test that spies on the system message sent to the LLM and asserts the Chinese example is gone and the language rule is anchored on the USER text.
Collaborator
Author
🤖 Open Code ReviewTarget: PR #2471 ✅ OpenCodeReview: Review complete: 0 finding(s) across 1 selected item(s). Generated by cloud-assistant via Open Code Review. |
Collaborator
Author
🔧 Open Code Review requested Agent fixOpen Code Review found 2 issue(s). I have resumed the development Agent to fix them.
The Agent will push a new commit to this PR branch. OCR will recheck after the commit is pushed. |
Address OCR findings on PR MemTensor#2471 without reintroducing the CJK anchor that originally caused issue MemTensor#2469. OCR finding 1 (summarizer.ts L106-L107): the single English example ("English text gets an English summary") may still anchor the model to English for non-English conversations. Add a French worked example and spell out that the rule applies to every other language. OCR finding 2 (summarizer.ts L110): the English-only "Do NOT prefix with 'The user said'" rule may let the model prepend an equivalent opener in a non-English conversation (e.g. a Chinese "用户说了…"). Rather than re-adding the specific CJK anchor — issue MemTensor#2469 proved that a lone "用户说了" example drove 89% CJK output on tool-call summaries with openai/gpt-6-luna, which the removal fixed to 0/30 — make the rule explicitly language-agnostic ("the equivalent 'the user said…' opener in any other language"). That covers Chinese, Japanese, Korean, French, etc. without priming any one of them. Extend the regression test with two new cases: one asserts the positive rule carries both English and French worked examples, the other asserts the negative-prefix rule reaches beyond English via an "any other language" clause. The two original guards (no "用户说了", no Han characters at all) still hold.
Collaborator
Author
❌ Automated Test Results: FAILED
Error detailsBranch: |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixed issue #2469: the memos-local-plugin capture summarizer was emitting Chinese summaries for English conversations on some models (e.g. openai/gpt-6-luna at temp 0, measured ~89% CJK on tool-call summaries and ~66% on conversational summaries). Root cause: the SYSTEM_PROMPT in apps/memos-local-plugin/core/capture/summarizer.ts anchored its language rule on an ambiguous referent ("in the user's original language") and carried a lone CJK example ("用户说了") that primed models toward Chinese under the 100-character cap. gemini-2.5-flash-lite was 0/750, confirming the leak only triggered on certain models.
Applied the reporter-validated diff: replaced the language rule with "written in the same language as the USER text (English text gets an English summary)" and removed the Chinese example from the "Do NOT prefix" rule. Added a regression unit test at tests/unit/capture/summarizer-prompt.test.ts that spies on the system message the summarizer sends and asserts (a) no Han characters leak into the prompt and (b) the USER-text anchoring is preserved — guarding both failure modes against future drift.
Verification: ran the regression tests against the pre-fix prompt first (both failed as expected per TDD), then applied the fix and re-ran. Full plugin unit suite is green — 182 test files / 1592 passed / 1 skipped. TypeScript check (tsc --noEmit -p tsconfig.json) exits 0. Committed on bugfix/autodev-2469-20261008201723155 and pushed to origin; opsp task file archived to the sibling specs repo.
Related Issue (Required): Fixes #2469
Type of change
Please delete options that are not relevant.
How Has This Been Tested?
Automated tests are pending.
Checklist
@whipser030, @hijzy please review this PR.
Reviewer Checklist