Add simple app & query functionality#1
Merged
Merged
Conversation
This commit enables anyone to create a app and add 3 types of data sources: * pdf file * youtube video * website It exposes a function called query which first gets similar docs from vector db and then passes it to LLM to get the final answer.
Adds a base chunker from which any chunker can inherit. Existing chunkers are refactored to inherit from this base chunker.
raghavtyagii
pushed a commit
to raghavtyagii/embedchain
that referenced
this pull request
Sep 25, 2023
merlinfrombelgium
pushed a commit
to merlinfrombelgium/mem0
that referenced
this pull request
Jul 4, 2025
Add simple app & query functionality
xiangnuans
added a commit
to xiangnuans/mem0
that referenced
this pull request
Dec 5, 2025
add openrouter and deepinfra llm
ywmail
added a commit
to ywmail/mem0
that referenced
this pull request
Jan 19, 2026
[TicketNo.] US20251217073371 mem0ai#1 [Binary Source] NA
This was referenced Feb 17, 2026
7 tasks
utkarsh240799
added a commit
that referenced
this pull request
Mar 13, 2026
- Add user identity to extraction preamble so memories are attributed to the correct user instead of cross-referencing cached patterns (OPE-6 #1) - Skip mem0.add() when no user messages remain after noise filtering, avoiding wasted API calls on assistant-only payloads (OPE-6 #2) - Raise auto-recall threshold to 0.6 (vs 0.5 for explicit search) and add dynamic thresholding that drops memories below 50% of the top result's score to reduce irrelevant context injection (OPE-6 #3) Co-Authored-By: Claude Opus 4.6 <[email protected]>
utkarsh240799
added a commit
that referenced
this pull request
Mar 16, 2026
- Add user identity to extraction preamble so memories are attributed to the correct user instead of cross-referencing cached patterns (OPE-6 #1) - Skip mem0.add() when no user messages remain after noise filtering, avoiding wasted API calls on assistant-only payloads (OPE-6 #2) - Raise auto-recall threshold to 0.6 (vs 0.5 for explicit search) and add dynamic thresholding that drops memories below 50% of the top result's score to reduce irrelevant context injection (OPE-6 #3) Co-Authored-By: Claude Opus 4.6 <[email protected]>
utkarsh240799
added a commit
that referenced
this pull request
Mar 16, 2026
- Add user identity to extraction preamble so memories are attributed to the correct user instead of cross-referencing cached patterns (OPE-6 #1) - Skip mem0.add() when no user messages remain after noise filtering, avoiding wasted API calls on assistant-only payloads (OPE-6 #2) - Raise auto-recall threshold to 0.6 (vs 0.5 for explicit search) and add dynamic thresholding that drops memories below 50% of the top result's score to reduce irrelevant context injection (OPE-6 #3) Co-Authored-By: Claude Opus 4.6 <[email protected]>
VictorECDSA
added a commit
to VictorECDSA/mem0
that referenced
this pull request
Mar 23, 2026
…operations # Fixes mem0ai#4490: Preserve original `actor_id` during memory UPDATE operations Fixes mem0ai#4490 ## Summary This PR attempts to fix an issue where `actor_id` metadata gets overwritten during UPDATE operations, breaking actor-level memory isolation in multi-actor scenarios. I've tested this fix locally and it seems to resolve the problem, but I would greatly appreciate maintainer review to ensure it aligns with the project's design goals. **Changes**: - Modified `_update_memory()` in `mem0/memory/main.py` (lines 1251-1252 and 2347-2349) - Changed conditional preservation to unconditional preservation of original `actor_id` ## Problem Statement Please see issue mem0ai#4490 for full details on the problem. **Quick summary**: When using `metadata={"actor_id": ...}` with shared `user_id` in multi-actor scenarios, the `actor_id` field appears to be overwritten when a different actor triggers an UPDATE event: ```python # Actor A creates memory m.add([{"role": "user", "content": "I am player mem0ai#1"}], user_id="team", metadata={"actor_id": "Alice"}) # ✓ Memory created with actor_id="Alice" # Actor B updates A's memory m.add([{"role": "user", "content": "Player mem0ai#1 is a good person"}], user_id="team", metadata={"actor_id": "Bob"}) # ❌ Memory updated but actor_id becomes "Bob" # Query fails m.search(query="", filters={"actor_id": "Alice"}) # Returns empty! ``` **Impact**: 1. Cannot query by original creator after UPDATE 2. Memory ownership tracking lost 3. Actor isolation fails in shared `user_id` scenarios ## Changes Made ### Code Changes **File**: `mem0/memory/main.py` **Location**: - Lines 1251-1252 (sync version of `_update_memory`) - Lines 2347-2348 (async version of `_update_memory`) **Before**: ```python if "actor_id" not in new_metadata and "actor_id" in existing_memory.payload: new_metadata["actor_id"] = existing_memory.payload["actor_id"] ``` **After**: ```python if "actor_id" in existing_memory.payload: new_metadata["actor_id"] = existing_memory.payload["actor_id"] ``` **Rationale**: Based on my investigation, it appears the condition check always evaluates to `False` because `new_metadata` contains the current actor's ID (passed from line 1668/1662). Removing this condition seems to make `actor_id` preservation work as (I believe) intended. This change treats `actor_id` as **"memory owner"** rather than **"last updater"**. The history table already tracks all contributors via `db.add_history()`, so update history is preserved. ## Testing ### Test Methodology I've attempted to validate this fix using a monkey patch approach that applies the proposed change to my local codebase without modifying source files. This allowed me to test the behavior before submitting the PR. ### Before Fix (Bug Demonstration) **Test scenario**: Alice creates a memory, Bob updates it with different `actor_id` **Result**: ``` === State AFTER Bob's UPDATE === Event: UPDATE This actor_id: 'Bob' ← Overwritten! ISOLATION TEST: Query Alice's memories: 0 found ❌ BUG: Alice's memory lost! ``` ### After Fix (Monkey Patch Validation) **Applied patch**: ```python from mem0.memory.main import Memory as MemoryClass def _patched_update_memory(self, memory_id, data, existing_embeddings, metadata=None): # ... (identical to original except actor_id handling) # ===== FIX ===== if "actor_id" in existing_memory.payload: new_metadata["actor_id"] = existing_memory.payload["actor_id"] # =============== # ... (rest unchanged) MemoryClass._update_memory = _patched_update_memory ``` **Result**: ``` === State AFTER Bob's UPDATE === Event: UPDATE This actor_id: 'Alice' ← Preserved! ISOLATION TEST: Query Alice's memories: 1 found ✅ SUCCESS: Alice's memory preserved! ``` ### Key Verification Points ✅ UPDATE event correctly triggered (not ADD) ✅ Memory content correctly updated ✅ `actor_id` preserved as original creator ✅ Query filtering by `actor_id` works correctly ✅ No side effects on other metadata fields ## Backward Compatibility To the best of my understanding, this change should not introduce breaking changes: - Existing code should continue to work without modification - API signature remains unchanged - Only affects behavior in multi-actor UPDATE scenarios - History table continues to track all contributors However, I may be missing some edge cases, so maintainer review would be greatly appreciated! ### Behavior Change Details **Before this fix:** When calling `add()` with a different `actor_id` that triggers an UPDATE event, the `actor_id` would be overwritten to the new actor. **After this fix:** The original `actor_id` (memory creator) is preserved during UPDATE operations. **Scenarios NOT affected:** - ✅ Calling `update()` method directly (already preserves `actor_id` correctly) - ✅ Updating memories without `actor_id` field (behavior unchanged) - ✅ Adding new memories (no UPDATE event triggered) **If you were relying on UPDATE to change `actor_id`:** You should use this pattern instead: ```python # Delete the old memory m.delete(memory_id) # Add a new memory with new actor m.add(..., metadata={"actor_id": "new_actor"}) ``` This change aligns with the semantic meaning of `actor_id` as the **memory creator** (immutable), not the **last updater** (tracked in history table). ## Additional Context ### Design Consistency This fix aligns with how other session identifiers are handled: ```python # Lines 1245-1250 already preserve these unconditionally (if not in new_metadata): if "user_id" not in new_metadata and "user_id" in existing_memory.payload: new_metadata["user_id"] = existing_memory.payload["user_id"] if "agent_id" not in new_metadata and "agent_id" in existing_memory.payload: new_metadata["agent_id"] = existing_memory.payload["agent_id"] if "run_id" not in new_metadata and "run_id" in existing_memory.payload: new_metadata["run_id"] = existing_memory.payload["run_id"] ``` The fix makes `actor_id` follow the same preservation pattern, but removes the always-false condition. ### Alternative Approaches I Considered I also thought about these alternatives, but decided against them (open to feedback though!): 1. **Track both `created_by` and `updated_by`** - Would require more complex schema changes - Less backward-compatible - Might not be necessary since history table already tracks updates 2. **Make preservation configurable** - Would add API complexity - I couldn't think of a clear use case for wanting `actor_id` to change - Seems to go against the semantic meaning of "actor" If either of these approaches seems better, I'm happy to revise the PR! ### Affected Implementations - ✅ Sync: `_update_memory()` (lines 1251-1252) - ✅ Async: `async _update_memory()` (lines 2347-2349) ## Checklist - [x] Code changes implemented (both sync and async) - [x] Tested with before/after comparison (monkey patch validation) - [x] Verified backward compatibility (to the best of my understanding) - [x] No breaking changes to existing API (as far as I can tell) - [x] Documented in PR description - [ ] Unit tests added (awaiting maintainer guidance on test location/structure) --- **Environment**: - Mem0 version: v1.0.7 - Python: 3.13 - Vector Store: Qdrant (local) Happy to make any adjustments based on maintainer feedback. Thanks for the great library!
11 tasks
8 tasks
Closed
kartik-mem0
added a commit
that referenced
this pull request
Jun 5, 2026
- require_admin: handle ADMIN_API_KEY and AUTH_DISABLED on fresh empty DB by falling through to allow bootstrap when no users exist yet, instead of raising 401 (Finding #1) - docker-compose healthcheck: use ${POSTGRES_USER:-postgres} instead of hardcoded -U postgres so custom POSTGRES_USER values work (Finding #2) - require_admin: restore request: Request param (needed for auth_type check) but now used in function body (Finding #7 resolved) - docs: sync migration guide password placeholder with README (Finding #9)
This was referenced Jun 8, 2026
8 tasks
14 tasks
This was referenced Jun 22, 2026
This was referenced Jul 2, 2026
18 tasks
14 tasks
14 tasks
14 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This enables anyone to create an app and add 3 types of data sources:
• pdf file
• youtube video
• website
It exposes a function called query which first gets similar docs from vector db and then passes it to LLM to get the final answer.