Rendered at 22:24:56 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
jrflo 7 hours ago [-]
Cool idea. Why is this beneficial over just using markdown files and allowing agents to grep for whatever they need? I've tried various MCP things in the past and I've found they tend to slow down the agent and waste tokens more than they end up helping, but a better memory system is 100% needed for agents.
schainks 3 hours ago [-]
At hundreds of notes you blow a lot of tokens just to _find one thing_.
Indexing your corpus as you go makes retrieval a lot faster, and then the agent can dig into the specific file if it needs something more.
cstrahan 5 hours ago [-]
The memories are stored as OKF (Open Knowledge Format), which is markdown + frontmatter (+ constraints/schema imposed thereon).
Having an inverted index (as with FTS5) is useful in that, for a basic single-term lookup, you reduce a sequential scan, O(N), down to O(log N). For small N, the performance difference might not be meaningful. Performance gap widens with more sophisticated queries (boolean operators, ranking, etc).
huflungdung 4 hours ago [-]
[dead]
agentifysh 6 hours ago [-]
im asking the same thing myself for personal projects seems markdown files is best.
i can see for public facing deployments agent memory like this could result in faster roundtrips.
rgbrgb 6 hours ago [-]
for one, the mcp-server architecture makes it usable from claude.ai and other surfaces where you have mcp but no filesystem. there are claude-specific workarounds (workspaces) but you lose portability across systems.
jrflo 6 hours ago [-]
Hmm, ok. I guess I rarely use the web interface and everything that I have agents record as "memory" in markdown is always accessible locally. If I'm accessing something remotely, I use the ChatGPT app with remote which connects directly to the host computer.
Natalia724 5 hours ago [-]
[dead]
infogulch 3 hours ago [-]
Last month a Show HN: ContextVault proposed an interesting long-term memory architecture. I discussed the design with the founder: https://news.ycombinator.com/item?id=48900288#48901679 (I think he bailed the conversation when I got too close haha.)
The basic shape is to periodically "distill the conversation into several areas (problem, solution, learnings, 'context' or original problem, plus other fields) and vectorized" (aka vector embedding), then queried against pgvector table to find related "memories". The vectorized distillates are also inserted into the pgvector table with a reference to back to the source conversation to add new memories.
Vector search requires a full scan but it's still pretty fast and I bet it's more accurate the FTS.
Alifatisk 4 hours ago [-]
What I do currently is having a markdown file named MEMORY.md at the root folder. Then, whenever I create a new conversation with an agent, I refer to that file. That file becomes the initial source of context and knowledge. At the end of an task, before I leave the conversation, I ask the agent to update MEMORY.md with lessons and new knowledge it has gathered from our conversation, it also removes stale or outdated information from that markdown file.
I myself do not care what's written in that file, I steer, instruct and share my knowledge, visions, goal and preferences in our conversations, and the agent will boil that down and update the markdown folder. It has worked very well for me.
I now do not have to worry about creating handoff prompts when creating a new conversation or that I have to teach an agent from the ground up about the context we're in, I just refer to that markdown file.
ksajadi 6 hours ago [-]
For those looking for similar tools, there is also https://markbase.cloud/ as a hosted service. (Disclaimer: we built it for internal use first and would like to open source with the community help as we don’t have much experience in OSS maintenance)
Since you built something on OKF, how would you contrast it with knowledge graph implementations? How do you manage the ontology of what to keep knowledge about? Any cases where traversal would have helped?
rcarmo 4 hours ago [-]
I’m not using pure OKF. Mine has links (both explicit and semantic, based on embeddings), and I tap into both Needle for routing searches and a Qwen/Gemma or gpt-mini for doing the actual traversals on behalf of the client.
dofm 8 hours ago [-]
Please excuse my noob-ish, naïve question, but to what extent is the business of getting the LLM to actually consult memory a model-dependent thing? Do you have to introduce the tool and guide models with different language for different model families?
Looking at your tool descriptions (as wit the ones on the original post) I wonder if this something perhaps only current frontier models will do, but the systems themselves seem like they'd be even more useful for open weights models with shorter working contexts.
rcarmo 4 hours ago [-]
All SOTA models seem to work fine (including Sonnet and Gemini), as does Kimi and DeepSeek. But this is not for short-term memory.
esafak 7 hours ago [-]
I have seen the value of recording past sessions but I am more skeptical of the value in recording facts, which may soon become stale, about a constantly changing code base. Got benchmarks?
rcarmo 4 hours ago [-]
This is not for code bases, that’s pointless. This is for durable facts like “this is the prod server” or “this is the skill for managing GitHub Actions cleanup policies”.
In short, this is for my agents to have a shared skill library, a shared fact library and durable information such as which projects run where.
The rest should be in your repo.
sho 7 hours ago [-]
Well, do you see the value of writing notes for yourself occasionally, even though they might soon become stale in your constantly changing environment? Yes, right?
Same principle. It's a good idea to have a schedule to clean them up periodically - an idea you can also put into a note.
esafak 4 hours ago [-]
I'll write one off notes for myself, but I am not going to do that in the code base unless it is really high value; it does not scale. The only place I write myself now is AGENTS.md
rcarmo 4 hours ago [-]
Memento is for that kind of cross-project, long term notes.
bravura 4 hours ago [-]
What if the memory were git repo backed, and the FTS5 were a speed-specific optimization?
Then the memories could easily be human-reviewed. The repo would be the canonical source, and the FTS5 would be one specific materialization.
ksajadi 3 hours ago [-]
that's what we started with, exactly because of the human review part. however we soon found out that there will be a lot of merge conflicts in the text (usually markdown) which heavily depends on the instructions and structure of the data. that's why for Markbase we opted for etag based validations to avoid any potential of conflicts.
clemens1010 7 hours ago [-]
did you test if that actually outperforms local claude code memory by any metric?
rgbrgb 6 hours ago [-]
that's a good idea. how might you test this? could also include a codex memory test.
I'm guessing having a portable memory that's comparable with first party memory is the goal.
healthycoder 6 hours ago [-]
How is this any different from all the other Memory stuff we have? mem0 etc etc that do the same thing?
rgbrgb 6 hours ago [-]
agree there are a lot of these but they're all pretty simple (including mine [0]) so I think building your own and playing around with architecture is useful and fun.
Looks very interesting. Can you explain for noobs why using Google's OKF format and not plain MD files?
pcbmaker20 7 hours ago [-]
OKF is basically md files with front-matter for meta data
bearjaws 7 hours ago [-]
Another week, another agent memory system that is about the same as grep in a memory/ directory.
cstrahan 5 hours ago [-]
I'm not sure I follow. Are you suggesting that full text search systems are ultimately a convoluted way of performing O(N) regex searches? If not, I don't see how you arrive at the conclusion that this is "the same as grep in a memory/ directory".
0c3ca83 6 hours ago [-]
How is this different from what's built into Claude?
myshapeprotocol 7 hours ago [-]
Using SQLite FTS5 for fast agent memory is such a pragmatic architectural choice. Great Show HN project.
Indexing your corpus as you go makes retrieval a lot faster, and then the agent can dig into the specific file if it needs something more.
Having an inverted index (as with FTS5) is useful in that, for a basic single-term lookup, you reduce a sequential scan, O(N), down to O(log N). For small N, the performance difference might not be meaningful. Performance gap widens with more sophisticated queries (boolean operators, ranking, etc).
i can see for public facing deployments agent memory like this could result in faster roundtrips.
The basic shape is to periodically "distill the conversation into several areas (problem, solution, learnings, 'context' or original problem, plus other fields) and vectorized" (aka vector embedding), then queried against pgvector table to find related "memories". The vectorized distillates are also inserted into the pgvector table with a reference to back to the source conversation to add new memories.
Vector search requires a full scan but it's still pretty fast and I bet it's more accurate the FTS.
I myself do not care what's written in that file, I steer, instruct and share my knowledge, visions, goal and preferences in our conversations, and the agent will boil that down and update the markdown folder. It has worked very well for me.
I now do not have to worry about creating handoff prompts when creating a new conversation or that I have to teach an agent from the ground up about the context we're in, I just refer to that markdown file.
Looking at your tool descriptions (as wit the ones on the original post) I wonder if this something perhaps only current frontier models will do, but the systems themselves seem like they'd be even more useful for open weights models with shorter working contexts.
In short, this is for my agents to have a shared skill library, a shared fact library and durable information such as which projects run where.
The rest should be in your repo.
Same principle. It's a good idea to have a schedule to clean them up periodically - an idea you can also put into a note.
Then the memories could easily be human-reviewed. The repo would be the canonical source, and the FTS5 would be one specific materialization.
I'm guessing having a portable memory that's comparable with first party memory is the goal.
[0]: https://setoku.com