Most agent frameworks coordinate calls to a model. The application manages the workflow, while the model s live state sits behind an API. We spent a year building around a different idea: make attention state something application code can program.
In Lloyal, the harness and model run in the same process. Your TypeScript code can fork the KV cache (the model s working memory) into concurrent agents.
Imagine the model has read a set of documents. You fork that state into agents investigating different explanations. They inherit what it has already read, then develop their own reasoning paths. The shared context stays shared; each branch accumulates its own work.
Your application decides what evidence reaches each branch, which tools it can use, when to stop it and which findings should survive at inference-time. It can discard an unproductive branch and reclaim its memory while the others continue.
Lloyal
Hello Product Hunters!
I’m Zuhair, founder of Lloyal Labs. Today we’re launching lloyal-ai: a platform to turn open-weight models into real AI apps people can download and use.
AI developers have two choices: rent intelligence from external providers, or assemble a stack with an inference server, an orchestration framework, a vector database and Docker containers.
Assembly is the easy part. Shipping that stack as a single app users can download and expect to “just work”? Not so easy.
Then there’s the deployment question: Can the same AI application run offline on a laptop for one person, serve multiple users from an on-prem appliance like NVIDIA DGX Station, and scale to frontier GPU clusters?
In the age of API wrappers, we took the bold step of deleting the HTTP boundary between the model and the harness, bringing them into one process.
Here's what emerged from that:
1. Programming live attention state: Your TypeScript code can fork the KV cache into concurrent agents on the GPU. They inherit what the model has already read, pursue different reasoning paths, and your application governs the convergence.
2. Composition of models: Give reasoning agents resident specialists for retrieval, ranking, vision out of the box or add your own specialists for classification, recommendations and other use-cases. Compose them as services in ordinary TypeScript or expose expert models as tools for the reasoning model.
3. Tools that participate in live inference: Tools in lloyal have superpowers: they access agent state, spawn sub-agents and respond to what the cohort has already done. Their results feed directly into the same live context.
And the form factor question: The same app runs offline on your laptop under Desktop target, can be served to multiple users and scale to frontier hardware under Web target.
Under the hood: Shared agent context along with liblloyal's Continous Tree Batching algorithm (https://github.com/lloyal-ai/lib...) means concurrent agents are paid for in space rather than compute (RunPod's CEO Zhen Lu posted about this: https://www.linkedin.com/posts/z...)
Get started:
npx lloyal-ai new
No API key. No Docker, Ollama, LangGraph, Mastra or vector database required.
Just your code and the model - what will you build?
Lloyal
Kazim here, the other maker. The best way to see what Zuhair's describing is to scaffold an app and keep the dev pane open while it works:
Pick a template (we recommend research) and keep the desktop surface ticked. Then:
First launch fetches and verifies three weights: a 4B reasoning model, a 0.6B reranker and a vision projector. Then ask it something worth investigating. You'll see the planner write an outline, a shared spine fork off it, and research agents inherit that spine and search at the same time. One model, inside your app's process, no API key.
Everything it generated is your TypeScript to change: the interface, the prompts, the tools, how the agents work. When you're happy, npx lloyal-ai ship --notarize builds a signed Mac installer (bring your Apple credentials).
Two minutes, start to finish: https://youtu.be/gp6LbadjSIY
If you try it, tell us where it got stuck. And what would you build on it?
@zuhairnaqvi @kazim_musa_ll The no api key part is pretty interesting. Congrats on the launch! curious to see where you take this
Lloyal
@zuhairnaqvi @harshyadav thank you! Would like for you build something, we are here if have any questions!