Graphs and Context
Categories AI Shit

Graphs and Context

Human context (understanding the big picture around a topic) and AI context (a llm’s current working memory) are two similar things that need to be managed differently. As humans we sort things relentlessly in our brains. We categorize and associate objects with systems. When we think of a bonsai tree we immediately touch japan, trees, calm, etc. We do not think of the next most likely word according to the entire corpus of human written knowledge.

One other observation that I’ve pondered lately: Walls of text are a very rich data source to an LLM. A picture is worth a thousand words to a human, but for llm’s, a thousand words are worth a few thousand tokens. Pictures have to be translated into words and tokens to be understood and reasoned with. One place this has become painfully aparent is in llm’s output to humans. It thinks(?) we are like it, able to consume endless words on a page. The more the marrier. The reality is that we tire quickly of trying to turn those words back into the conceptual system that lives in our heads. We deal much better with high-bandwidth communication. A picture or conversation or diagram is much faster to digest. For example:

This diagram communicates with greater efficiency than another paragraph of text about how an llm’s processing is different than ours. I understand it faster, but it takes more effort to compile.

This brings us to my hypothesis. So far the best version of “memory” I’ve seen for llm’s is simply .md files in a repo. This works for a lot of simple use cases but falls apart when the library of information becomes too large. Too many MD files, not enough pruning of documentation, and multiple developers working on different things start to muddy the waters. Some degree of file structuring helps, but now what you have is a directory of such dense, unreadable, and maybe contradicting text that a human has no hope of reading it, let alone check it for accuracy. Our context is now splintered with the AI, and we have no choice but to accept the output at it’s word.

I need a way to store this information that is queryable and readable by a human AND an llm. I need provenance, the ability to discredit a fact if found to be contradictory. I need a graph database.

Relational databases are great for storing a lot of items that all have the same conceptual shape. Graph databases so far have been good at doing that same thing, but being more descriptive about the relationships. Edges add context that relational databases lack, properties allow for variations on each node’s data. The query language allows for much more expressive and natural language like queries. I love Bryon Jacob’s thoughts here on why RDF is the natural fit for AI systems. His whole series there around RDF* and OWL ontologies has been influential over the last year of thought.

So if Graph databases and triples are the natural fit, which engine should we use? Again, a cliff hanger, but spoiler alert: none of the mainstream db engines right now are well suited for agentic knowlege base workloads. There are three things I was looking for that seemed to be lacking.

  1. Bitemporal Data model support
  2. Support for hot-loading/dynamic ontologies
  3. HNSW vector search alongside the graph data.

More on that next time.

Prev Shared Human-AI Context

Leave a Reply

Your email address will not be published. Required fields are marked *