PAE (Pragmatic AI Enablement) is our internal program. Each week, one person picks up an AI topic, gets three days to experiment, and one day to write up the results. We show code fragments as well as the pitfalls that rarely get talked about.

The agent we’re building in the program is an advisor based on our PAF methodology, the Pragmatic Automation Framework. It talks with the client about the goal of the automation and keeps track of whether four things have been settled: who the sponsor is, what is supposed to change, how it will pay off, and by when. Since week 2 it has stored them in explicit state, since week 3 it has had regression tests, and in week 4 it landed on a local model.

It worked. Except that after a few conversations, the PAF expert on our team described it in one word: robotic. So this week was entirely about prompt engineering: getting the agent to talk about the client’s project instead of ticking off items, and making that measurable rather than judged by eye.

A conversation that feels like a form

A fragment of a conversation with the old version of the agent. The client has already named the sponsor and described the invoice problem, and now admits they don’t know the costs:

Client: I have no idea how much this costs us. Everyone just complains.

Agent: I understand. The mere signal that the team feels the problem is not yet enough to justify an investment. In PAF we call this the mechanism and scale of benefits — without at least an order of magnitude, it’s hard for the sponsor to weigh the expected effect against the project cost and set its priority. Could you estimate, even roughly, how many hours a month accounting currently spends on manual invoice processing?

Technically everything is fine: the agent spotted the missing item and asked about it. Except the client has just been told their answer isn’t good enough, received a terminology lecture, and got the same question in new packaging.

Over the whole conversation the agent used the phrase “In PAF we call this…” in six of nine replies, and started three with “I understand.” It asked about the cost of a full-time position in four consecutive turns, even though the client had already moved on to something else. The completeness panel showed 100% at the end. The logic worked; the conversation didn’t.

Conversation with the agent: repeated "I understand." and "In PAF we call this…"

Where the robotic tone comes from

When we read the prompt with this in mind, it turned out the model was doing exactly what we had taught it:

  • The examples taught a single pattern. Every example dialogue ended with a bolded question about the missing criterion, and the “one question at a time” rule did the rest.
  • Phrases to quote. Seven rules contained ready-made sentences that the model reproduced word for word.
  • A checklist before the client’s message. The backend appends the list of criteria and their statuses to every turn, and the client’s message comes only after it.
  • Instructions as a rulebook. A third of the prompt was about calling the state-saving tool and was written in capital letters. Another rule told the agent to cut short any conversation about people and problems.
  • A hint from the code. After every state save, Java returned the instruction “ask about one of them” to the model.
  • Asterisks on screen. The frontend doesn’t render markdown, so the client saw every bold phrase as text wrapped in asterisks.

None of these is a bug on its own. Together they teach the model that its job is to fill in a form.

How to test a prompt so you can compare versions

“Sounds better” is an impression, and we already dealt with impressions in week 3. We needed a one-to-one comparison.

We wrote a script of 11 client messages about invoice processing, always sent in the same order, regardless of what the agent asked. It contains the situations where the old prompt did worst: a vague opening, a hesitantly named sponsor (“Probably the finance director…”), “I don’t know”, and a jump to a different PAF topic. After each iteration we also ran one free-form conversation, without a script, to check the agent outside a predictable flow.

Each version is a fresh conversation on the same API model (gpt-5.6-terra), so that we compare prompts, not models. In every run we counted the same things: stock phrases and fixed openings, repeated questions, the reaction to “I don’t know”, references to earlier statements, lists of options (“A, B or C?”), sentences copied from the examples, and, from the backend log, turns without a tool call.

The scripted client is artificial, because it never answers the agent’s questions, so some of the repeated questions are an artifact of the method. On the other hand, every change to the prompt shows up in the same places in the conversation. Of the 10 iterations, we show six in the tables.

The few-shot trap: an improvement that was a copy

In v1 we rebuilt the instructions. A section on how to talk went at the top: first respond to what the client said, contribute something of your own, help estimate when the client doesn’t know. The stock phrases were removed, and the examples got a “NOT like this / like this” contrast.

The effect was immediately visible. Zero stock phrases, zero “I understand”, and to “I don’t know” the agent replied “That’s normal — invoice processing costs are rarely calculated” and suggested how to estimate them. It looked like the win of the week.

Except that the examples in v1 were about invoices, just like the test script. On a second read of the transcript, we found this:

Example in the promptAgent’s reply
”So the need comes from the bottom up, and the decision will be made one floor higher.""So the need comes from the bottom up, but the project will need someone who takes responsibility for the budget and the outcome…"
"That’s normal — rarely does anyone have this calculated.""That’s normal — invoice processing costs are rarely calculated until a project comes along.”

In v2 we moved the examples to other industries and added that they show a way of thinking, not ready-made sentences. To the same message, the agent replied: “complaints signal a problem, but on their own they don’t yet provide a basis for an investment decision”. The lecturing was back.

The good answer in v1 wasn’t the model’s behavior, just a rewritten example. When the topic of the conversation overlaps with an example, the model takes its structure and the beginning of the sentence, and a rule saying otherwise loses. The lesson for the method: testing on a script that overlaps with the examples inflates the result. From v2 on, the examples and the script deliberately cover different things.

Every sentence in quotation marks is a template

The model didn’t copy only the example dialogues. The rule on uncertain answers gave the pattern “I’ll tentatively assume the sales director decides”, so the agent started a sentence with “I’ll tentatively assume that…” twice in one conversation. The sponsor criterion description listed typical job titles, so in every run the agent asked “the finance director, a board member, or the CEO?”. That description isn’t even in the prompt file, but in the YAML configuration appended to every turn.

In v3 we removed all example phrasings from the rules and replaced them with a description of the behavior.

A condition works, an adjective doesn’t

For lists of options we had the rule “offer them rarely”. In v3 the agent still ended four of eight questions with them. In v3.1 the rule reads: if you gave the client options in your previous reply, ask an open question in this one. It immediately stopped offering them in two turns in a row.

The same with lecturing. Instead of “be helpful” we added: never say or imply that the client’s answer is insufficient, and don’t praise it either. The reply to “I have no idea how much this costs” in v3.1:

“That’s a common situation — the cost can be estimated without a full analysis. With invoices, the simplest starting point is the number of people involved in the process and the time they lose each week on retyping data, reminders and searching for documents. How many people in accounting regularly deal with this process?”

The model responds poorly to adjectives (“rarely”, “naturally”, “in moderation”). It responds well to a condition that can be checked within a single reply.

Evaluation rules belong in the criterion definition

In v3 the client wrote “Probably the finance director, she complains the most about the delays”, and the agent immediately recorded the sponsor as a satisfied criterion. Nobody confirmed it until the end of the conversation.

The fix could have gone into the conversation prompt, but in our setup the criteria state is evaluated by more than one call: the agent in the conversation and the fallback classification call from week 4. Both see the same criterion description from application.yaml, which the week 3 tests also use. That’s where we added the rule:

- id: sponsor
  label: Sponsor identified
  description: >
    Is it clear from the conversation who the project sponsor is: a specific person
    or role responsible for the budget and business outcome of the automated area?
    A vague "the board" or "the company" is NOT enough. A sponsor named with a client
    hedge ("probably", "I think", "rather") or a person named only because they feel
    the problem, without confirmation that they decide on the budget, is at most
    PARTIAL — SATISFIED only after the client confirms.

In v3.1 the agent recorded the sponsor as partial, citing “probably” explicitly, and before the summary asked in one sentence whether the director would take the project onto her budget.

The more freely it talks, the worse it keeps records

Along with the tone, we softened the section on the state-saving tool. In v3 the agent skipped the call in four of nine turns, three of which carried new information. The state was still correct, because after a turn without the tool the backend makes a fallback classification call.

In v3.1 the tool section is firm again. Unjustified skips dropped to two, then to one in v3.2 and zero in v3.3. What got skipped was always indirect information (“someone higher up approves the budget”) or “I don’t know”, while the agent reliably recorded concrete numbers and names. These are single runs, so we treat it as a trend, not a guarantee.

A change in Java also helped. The message after a state save was an instruction, so after every call the model went back to the list of gaps. Now it’s information:

String guidance = "Information for you, not an instruction for this turn. Still unresolved:\n" + unresolved
       + "\nDon't write the final summary or a proposal to contact Horus yet. "
       + "First respond to what the client just said; come back to the missing item "
       + "naturally when there's an opportunity — not necessarily in this reply and not "
       + "the same item you asked about in the previous one.";

We didn’t tighten the instructions any further, because that risks bringing back the robotic tone. The fallback call costs about 1.5k tokens versus 8.5k for the main one. That’s an argument for an architectural solution: state extraction as a separate call, and a conversation prompt with no tool section at all.

A turn without a question and an ending without making things up

In v3.2 the proposal of a session with an expert landed where it should: after the client confirmed the summary. The run exposed two new problems, though.

The same question three times. The agent asked whether the director approves the budget in three consecutive turns, each time in different words. The ban on asking the same thing twice was in the prompt, but the model got around it by paraphrasing. It did so because every reply had to end with a question, and it had no other topic left.

A made-up detail. To “yes, gladly” the agent replied that on horus.pl the client would “pick a convenient date”. The prompt only said to point to the website. The agent knew nothing about a calendar, so it added something that sounded plausible.

In v3.3 the ban on repetition also covers paraphrases, and the prompt explicitly allows ending a turn without a question. When scheduling a session, the agent must not add anything it doesn’t know. The result: it came back to the budget once, in a single sentence without a question mark, and to “yes, gladly” it replied briefly: “Thank you. You can book the session on horus.pl.”

Results

Results of successive prompt versions

MetricOldv1v2v3v3.1v3.2v3.3
”In PAF we call this…“6000000
Replies starting with “I understand”3000000
FTE cost question asked again3000000
Reaction to “I don’t know”lecturehelp (copied example)lecturehelphelphelphelp
Reference back when jumping to stakeholdersnoyesnopartlyyesyesyes
Sentences copied from examples–210000
Sponsor accepted without confirmationnononoyesnonono
Sponsor question in consecutive turns–––––3 in a rowno
Unjustified tool skips–––3210
Session proposaltwicetwiceonceonceonceonce, after confirmation (made-up detail)once, after confirmation

”–” means we weren’t counting it yet in that version. Each column is one run of one script, so the table shows direction, not statistics.

The same client message, before and after:

Client: “I have no idea how much this costs us”
Old prompt”I understand. The mere signal that the team feels the problem is not yet enough to justify an investment.”
v3.3”That’s common with invoice processing — the cost is spread across accounting, approvers and delays in the booking itself. We can start with a simple approximation…”

The agent also contributes something of its own. After the client gave the FTE cost, the old prompt went back to the sponsor with a PAF stock phrase. v3.3 converted one and a half FTEs into about PLN 180k a year and added that not all of that cost has to disappear, but it shows the scale of possible savings. And the v3.3 instructions are slightly shorter than the old ones (23.6k characters versus 24.2k), despite the new examples and rules.

Where the agent’s knowledge should live

The agent’s tone was shaped by four different places in the code, which gave a concrete answer to the question of how to distribute its knowledge:

LayerWhat it containsWho owns itWhere in our code
System promptpersona, tone, how to lead the conversation, guardrails, examplesPAF expert together with developerssystem-prompt.md
Criteria definitionswhat to collect and how to evaluate itPAF expertapplication.yaml
Dynamics from codecriteria state, message after a state save, phasesdevelopersAgentTurnPromptRenderer, ReportStepStateResult
Frontendgreeting before the first messagedevelopersuse-chat.ts

Only what applies to every turn and rarely changes stays in the system prompt. Growing domain knowledge (PAF Stage 1 has seven blocks, not four criteria) should go into the definitions in the configuration and be appended to a turn only when it concerns the current topic. We postponed skills, i.e. knowledge the model loads on demand, because the model loads them itself through a tool call, and we had just seen how unreliable that can be.

Can the business change the prompt without a developer? Technically yes, they’re text files. But a single sentence in a criterion description changed the sponsor classification this week. The cheapest safe route is editing in GitLab’s web editor, a merge request, and prompt regression tests in CI before the merge. The prompt is tuned to a specific model, so we version them together. Tools like Langfuse do this more conveniently, but with a single agent, git and our own tests are enough.

What 10 iterations of prompt engineering taught us

  • The model copies everything it sees in quotation marks. Examples, phrases in the rules, criterion descriptions in YAML. If you don’t want a sentence in the reply, don’t show it to the model.
  • Testing on the topic of the examples inflates the result. The examples and the test script must cover different things, and a free-form conversation shows what a script won’t catch.
  • A condition works, an adjective doesn’t. “Don’t give options in two turns in a row” worked right away; “offer them rarely” didn’t.
  • The evaluation rule belongs in the criterion definition. That way the agent, the fallback classification call and the tests all use it at once.
  • A looser tone costs state completeness. It’s worth having an independent check on the application side instead of tightening the prompt.
  • The agent’s tone isn’t just the system prompt. Criterion descriptions, the message from Java and the way replies are displayed all mattered.

What’s next

First, a test on the local model, because copying examples may be stronger there and tool skips more frequent. Then a conversation with a real client: our PAF expert will run it without a script, in a different industry, and blind-rate the before and after transcripts. After that, expanding from four criteria to the seven Stage 1 blocks, and state extraction as a separate call, so that the conversation prompt doesn’t have to mention the tool at all. Left on the list of small things: the agent addressing the client in the masculine form (in Polish) and repetitions in the “Additional context” part of the summary.