What shifts about AI transparency when the harness and model share one runtime?

A model has read your documents. Three agents fork that working memory and pursue different reasoning paths. One calls a tool and takes in new evidence while the others keep generating.


What does that look like under the hood?


This is what we’re making visible with Lloyal’s developer tools. The model and harness run in the same process, and agents are branches of its live attention state.


The timeline shows their shared context and divergent execution. Inside a tool call, you can inspect which passages were selected to enter an agent’s context.

Alongside it, entropy and surprisal move as the model generates tokens. Individual agents in a cohort can be stopped while the others continue.

That’s the programming surface we’ve been building around: application code can fork, extend and prune the model’s working memory, not transcript but tensors. Lloyal dev-tools make that live execution inspectable and interactive.

Try it out by creating a local AI app today with

And share with us what total transparency AI feels like!

Concurrent agents and live model state during inference:


Entailment based context admission (JEV style decision on what to read from a fetched web-page):

8 views

Add a comment

Replies

Be the first to comment