What if spawning an AI agent meant forking the model’s working memory?

Most agent frameworks coordinate calls to a model. The application manages the workflow, while the model’s live state sits behind an API. We spent a year building around a different idea: make attention state something application code can program.

In Lloyal, the harness and model run in the same process. Your TypeScript code can fork the KV cache (the model’s working memory) into concurrent agents.

Imagine the model has read a set of documents. You fork that state into agents investigating different explanations. They inherit what it has already read, then develop their own reasoning paths. The shared context stays shared; each branch accumulates its own work.


Your application decides what evidence reaches each branch, which tools it can use, when to stop it and which findings should survive at inference-time. It can discard an unproductive branch and reclaim its memory while the others continue.


That changes what an agent is in your code. It has an addressable position in live inference, a parent and a lifetime you control.

Runpod’s CEO, Zhen Lu, : ten agents sharing a 239 GB GLM-5.2 model on two B200s for 82 minutes, costing approximately $16 in total GPU rental. He called it “the cleanest counterexample I've seen” to multi-agent reasoning multiplying inference bills.


The same harness runs a 4B model on a laptop.


Underneath, agents share the existing KV prefix and advance together through batched decoding. Memory capacity and growing context still impose real limits, so deciding when to admit work, stop branches and retain findings is part of programming the application.


I wrote up the architecture, execution path and a failure case we encountered here:


You can try it with:

npx lloyal-ai new


For people building with agents: what do you most need to control during a run that your current framework only lets you check afterwards?

23 views

Add a comment

Replies

Best

If you do end up trying it out, please leave your review in