From reactive agents to more autonomous ones: the missing piece is the state of work
Summary A reactive agent waits for an instruction and for context handed to it. A more autonomous agent has to know what has already been done, what is still expected and how far it may act. That gap is not a model problem. It is a state problem.
Two words worth separating first
“Agent” is used for two very different things. Before going further, here is what each one means.
A reactive agent waits to be asked. You give it the instruction, you give it the context, such as a file, a snippet of conversation or a link, and it produces a result. Between two requests it knows nothing. That is not a flaw; that is how it works.
A more autonomous agent chains several steps without being handed the context each time. It can decide to wait, to resume, to follow up, to stop. To do that, it needs a representation of the situation that survives between two runs.
So the difference is not, first of all, a difference in model power. It is a difference in memory of the situation.
Hand-assembled context has a natural ceiling
As long as a human assembles the context for every request, the agent inherits the quality of that assembly. This works well for an isolated task: writing, summarising, comparing, rephrasing.
It holds far less well as soon as the task stretches over time and crosses several people. The reason is simple: between two requests, the work carried on elsewhere. A decision was made in a meeting, a deadline moved inside a tool, a trade-off arrived by email. The agent picks up from whatever it was handed again, which means starting from a snapshot that is already out of date.
The cost is not only an imprecise answer. It is that the rebuilding work lands back on the human: reopening the threads, finding what changed, reassembling the context before being able to delegate again.
Three levels systems often blur together
Most tools know one of these three levels well, and rarely all three.
| Level | What it is | What carries it today |
|---|---|---|
| Expected work | What was meant to happen: tasks, commitments, deadlines, deliverables | To-do lists, project tools, plans |
| Observed work | What actually happened: meetings, messages, files, actions inside tools | Traces scattered across each tool |
| Current state | What moved, what changed, what is still open, and on what evidence | Usually: somebody’s head |
A project tool mostly describes the first level. A tracking system mostly describes the second. The third, which an agent actually needs in order to act without asking for everything again, is almost never maintained anywhere.
What “being able to act” demands, beyond knowing
Giving an agent the ability to act inside tools moves the question. It is no longer only about knowing what is true, but about knowing how far it may commit.
Four things then become necessary:
- A sourced state: every claim points back to what produced it, and stays contestable.
- An explicit mandate: what the agent may do alone, what requires validation, what is off limits.
- A human resumption point: being able to understand what was done and take back control without rebuilding the history.
- A usable trace: not to monitor people, but to explain a decision after the fact.
None of those four is solved by switching model.
The hypothesis I am testing
I build Private Twin to confront one precise hypothesis: if expected work and the signals of observed work are brought together, a current state can be maintained automatically and remain reliable enough for a human, and for an agent, to resume the work without rebuilding its history.
That is a hypothesis, not a demonstrated result. What exists today: a complete prototype and an evaluation bench that replays frozen days to catch regressions. What remains open: evaluation across more real situations, and validation of the value users actually perceive.
What this changes in practice
For an organisation deploying agents today, the practical consequence is easy to state: the question is not only “which agent should we build”, but “what state does this work need to keep so that a person or a system can pick it back up”.
An agent that performs in a demo on an isolated case proves a capability. An agent that picks a file back up three days later, without being re-briefed on where things stand, proves an infrastructure.