The agent from week 1 could hold a conversation, but had no idea when to end it — it didn’t know whether the client had said enough to consider the PAF task “Goal Agreement” closed. Week 2 was meant to change that: give the agent explicit conversation state, maintained across turns, and show it live as a completeness percentage.
The problem
Stage 1 of the PAF methodology (“Benefits”) requires four findings: sponsor, expected outcome, mechanism and scale of the benefit, and time horizon — three of them map directly onto the SMART criteria (S/M/T), while the remaining two letters (A/R) belong to other PAF tasks. In this iteration we deliberately left them out of scope. Without explicit state we saw two risks: the agent could close the conversation prematurely, before everything was established, or it could keep asking endlessly about something the client had already said.
The goal of this iteration was an endpoint and a panel that, after every turn, shows how many of these four
criteria are already met — without a separate completeness-assessment step by the model. We adopted a
deliberately strict rule: every criterion with status SATISFIED contributes 25 percentage points, and the rest
contribute 0. So we don’t treat this score as the “model’s confidence,” but as a plain answer to the question: how
many of the required criteria have we already satisfied.

Conversation state instead of asking the model every time
In week 1 we announced completeness assessment via structured output — a
response in which the model returns not just free text for the client, but also data in a strictly defined
format. The plan was simple: in one result, the model would return both the text and the conversation state. In
Spring AI, such state is retrieved via .entity(CompletenessJudgement.class): the framework takes the JSON from
the model’s response and maps it onto a CompletenessJudgement Java object.
This idea, however, runs into a streaming limitation. .entity() only works with .call(), which returns the
whole response once it’s complete, not with .stream(), which sends text in chunks. The model would have to
finish its entire response before the framework could read the state; the user wouldn’t see it token by token.
You can stream the JSON manually and parse it once it’s fully collected, but the state would still only be
available at the very end, and handling the stream on the UI side would be more complex.
So in this iteration we chose tool calling — the model uses it to report which parts of the goal it considers
established, before it even writes the reply for the client. A tool call is a request from the model asking the
application to execute a function it has been given access to; here, that function is reportStepStateTool,
which records the conversation-state assessment. The application executes it and, based on the stored state,
calculates the completeness percentage and whether the conversation can move on to confirmation. The panel only
reads that stored state.
The roles here are clearly separated. We leave content assessment to the model — it interprets what the client says and decides whether a given criterion is already met. The code doesn’t verify that. What the code does enforce is one rigid process rule: confirming the goal requires all four findings at once.
This isn’t the only way to track conversation state — we also considered, among others, a single call combining state and response, and a separate extractor model running alongside the conversational model.
A dead end: tool calling configured but never invoked
The completeness panel stubbornly showed 0%, regardless of how the conversation went — even though the
end-to-end mechanism looked correct: the reportStepStateTool tool was registered, the application calculated
completeness, the endpoint served it.
The exchange log between the application and the model settled it within a minute: in the response text, the
model correctly referred to the established sponsor, but across 12 turns and several sessions there wasn’t a
single tool call. From this we concluded that it was missing an explicit instruction to use the function in the
system prompt. In our configuration, the description of the reportStepStateTool function itself told the model
what it does, but not when to call it. So we added that second message directly to the system prompt — the
instruction appended to every turn of the conversation.
We found this manually, by going through the log — exactly the kind of check that week 3’s evaluations are meant to eventually take over.
State says one thing, text says another
The same mechanism also revealed a different kind of failure — the most interesting one for us this week. In a
conversation about automating order dispatch, the model correctly reported expectedOutcome as PARTIAL. Two
turns later, even though the client hadn’t said anything new about that criterion, the model wrote in its summary
that “the goal is agreed, all criteria are met” — and added a closing line inviting the client to get in touch
with Horus. The completeness panel, read deterministically from the state, without a second model call, showed
75% in that very same turn.
From a mechanical standpoint, the tool call worked as expected: it returned the correctly computed phase
(IN_PROGRESS) to the model — i.e., information that at least one criterion still needed to be established. All
four criteria being met produces the READY_FOR_CONFIRMATION phase; only an explicit confirmation from the
client changes it to CONFIRMED. The tool call’s result landed in the context right before further generation,
and yet the model wrote a narrative that contradicted it.
We implemented a first layer of defense: alongside the completeness assessment (phase), the
reportStepStateTool result now also carries a new guidance field with an additional instruction for the
model. In the NOT_STARTED and IN_PROGRESS phases — that is, as long as not all criteria are met — guidance
explicitly lists the missing criteria and forbids writing a summary. In READY_FOR_CONFIRMATION, CONFIRMED, and
BLOCKED the agent shouldn’t be asking further questions, so the tool returns different content:

With technical identifiers stripped out, the instruction given to the model looks like this when two of the four
criteria are still not SATISFIED:

This instruction reaches the model right before further generation, not just once, in a static prompt. In the literature, the phenomenon of important information getting skipped when it’s buried in the middle of a long context is sometimes called “lost in the middle.” We’re assuming that placing this instruction close to the decision point increases the chance the model will take it into account. We don’t, however, treat this as a guarantee.
This change is meant to reduce the risk of drift, but it offers no guarantee: the instruction is still only part of the model’s context. If we want a piece of the response to be strictly dependent on the state, the application should assemble it itself or check it after generation. For now we’ve implemented the first, prompt-based layer; the remaining solutions stay in the backlog.
This is a more important takeaway for us than the prompt fix itself: a prompt can lower the probability of an error, but it shouldn’t be the only mechanism enforcing a process rule. When the content of a message depends on state that’s critical to the process, the application should ultimately either check that dependency or build that part of the response itself.
A limitation observed along the way: the model running on our infrastructure (gemma4:26b) sometimes returned
corrupted data in a function call — invalid JSON, for example. In our setup we couldn’t assume the data returned
by the model was correct, so the application has to validate it on receipt and reject anything malformed. That’s
a separate axis: whether a locally run model produces good enough results. We’re planning a full comparison
against the API for week 5.
Result
| What works after this iteration | How it shows up |
|---|---|
| Explicit goal state | After every turn, the panel shows four criteria: sponsor, outcome, mechanism/scale, and time horizon. |
| Completeness percentage | 0%, 25%, 50%, 75%, or 100% — counting only fully met criteria. |
| Read without an extra LLM call | The panel and endpoint read the stored state; they don’t trigger a separate conversation assessment. |
| Process phases | Five phases from NOT_STARTED to CONFIRMED or BLOCKED; readiness requires all four criteria met. |
| Protection against state/text drift | The model receives a direct instruction derived from the current state; stronger enforcement stays in the backlog. |
From our perspective, the last row matters most: we have a working, readable state and an endpoint that exposes it, but the current protection against a contradictory narrative is only an instruction to the model. That’s a deliberately open question, not an oversight.
What’s next
The most important gap this week exposed for us wasn’t another guardrail, but a way of detecting errors: we found these ones manually, in individual conversations and logs. After the prompt change, we still don’t have an automated answer to the question of whether the agent still updates the goal state correctly, doesn’t end conversations too early, and doesn’t break behavior that used to work.
Week 3’s topic is evaluations and prompt regression tests: we’ll build a repeatable set of scenarios in Java that, after every change, shows whether the agent still behaves as intended.
Every prompt change is a hypothesis, not a fix: without measurement, all we know is that it worked in a handful of conversations. Building an agent is inherently iterative, and evaluations are what separates iteration from guessing — and only they let us responsibly wire the agent into a real business process.