How important is human-agent experience (HAX) in the AI space?
Memory (or context) is a very interesting problem in AI, but not for the reason most people think.
This is the first layer that needs to be EQUALLY READABLE BY MACHINES AND HUMANS (opinions are my own).
I have done an extensive analysis of the current products in the space, some of which are heavily funded like Mem0 or supermemory.
Current focus areas include making retrievals faster, hitting benchmarks, expanding the product surface (all valid problems).
But they're all missing the human-agent interaction part.
Right now, memory works like this: you give an context engine your memories as a dump and then it does its thing and start providing your agents the required context.
But keep doing it for a while (in any product of your choice) and then slowly it becomes a junk drawer where you know you put something but where was it, why did you add it, who changed it, poof. Good luck at that point.
This is one of the areas I'm obsessed with. How do we make context engineering as simple as a perfectly organized notion, and still ridiculously easy for agents to retrieve as well.
What I keep wrestling with in my mind these days: Imagine you are an AI-native individual with 100s of documents and dozens of ongoing clients, is a knowledge graph really the right abstraction?
What "context abstraction" would you want if you are working with agents day-in day-out?


Replies
A knowledge graph solves retrieval, not the junk drawer problem you are describing. The failure is not that memory is hard to search, it is that nobody knows whether a memory is still true. A graph will happily retrieve last month's client requirement three months after the client changed their mind, fast and structured, which makes it worse rather than better because structure makes it look authoritative.
What I would want is less abstraction and more expiry. Every memory gets a source and a date at write time, not just content. When two memories about the same thing disagree, the system should surface the conflict instead of quietly returning the newer one, because the newer one is not always correct, sometimes it is a mistake nobody has caught yet.
Running several products at once, my own junk drawer moment is always the same shape, a decision made for one client or product that got copied as a pattern into two others, and eighteen months later nobody remembers it was a one off exception rather than a rule. Structure would not have saved me there. A system that made me confirm the memory was still true would have.
Have you tried a freshness or confidence score on each memory instead of a graph relation, or does that just move the junk drawer into another field nobody fills in?
@oshylabs we literally have an internal brainstorming session in 3 hours on how to best portray freshness signals + conflict resolution.
I feel like we need something more than a knowledge graph, but i'm sure my team will bring in more ideas. Let's see.
At this point, AICF is the most user friendly AI context management system, but we are constantly making improvements in this direction.
What we have now: (source tagging, who added it, version history)
Breadcrumbs (another view that we are brainstorming on how to add) - my proposal for today's meeting:
A graph showing how the different buckets and their memories are connected. The brighter the dots, the more recent they are, the dimmer the dots, the older they are.
Health and conflict sidebars bring out the memories that are conflicting or stale, and you can auto-resolve them.
Our product philosophy is to keep this organization part manual i.e. more control for user - less chance for AI to do something with your memory that you did not intend.
Happy to have your honest feedback.
@hira_siddiqui1 The brightness by recency idea is the right instinct but the part that will actually get argued over in your meeting is what counts as recent. Timestamp of last edit and timestamp of last confirmed still true are two different clocks, and most systems only track the first one. A memory nobody has touched in six months looks stale by edit time even if it was reconfirmed correct last week by a completely different action, like a user re-selecting it in a form.
The conflict sidebar is the harder half. Surfacing a conflict is easy, deciding which side wins by default is where teams get stuck, because newest is not always correct, it is just newest. If two sources disagree and one is a user's direct statement while the other is an inference your system made from behaviour, those should not resolve the same way even if the inference is more recent.
Keeping resolution manual is the right call for exactly the reason you gave. The moment you let the system silently pick a winner, you have reintroduced the freshness problem one layer up, now hidden behind a decision nobody can see was made.