Policy Factorisation decomposes the agent's policy into reasoning and action (Wei et al., 2026).
LLM-based policies produce z in natural language. Is that a particular feature of Agentic AI?
{
"actor": "agent_3",
"kind": "agent_update",
"payload": {
"decision": {
"response": {
"choice": {"status": "act", "action": 1},
"reasoning": {
"confidence": 0.8,
"source": "llm",
"text": "Moving left is the least
explored direction...",
"tool_steps": [{"kind": "tool",
"name": "llm", "elapsed_s": 1.24}]
}
},
"explanation": "Moving left — least explored."
}
}
}
LLM record: action + observable reasoning (z).
{
"actor": "agent_3",
"kind": "agent_update",
"payload": {
"decision": {
"response": {
"choice": {
"status": "abstain",
"action": null
},
"reasoning": {
"confidence": 0.0,
"source": "llm",
"text": "All surrounding cells explored.
Cannot determine best move."
}
},
"explanation": "Abstained: low confidence."
}
}
}
Let's imagine the hypothetical case where the agent says "I Don't Know".
If we factorise the policy, DOAgent offers and engineering approach to observe the reasoning trace that explains why the agent abstained.
Deploying ML in the Intensive Care Unit to support practitioners in collaboration with clinicians at the Karolinska Institute.
Uniform access to harmonised ICU data across heterogeneous sources (aICU, MIMIC, eICU, HiRID, …).
from aicu_access import load_concept
df = load_concept("hr", source="miiv")
_script: true
This script will only execute in HTML slides
_script: true