Prompt Engineering in Practice: How We De-Roboticized an AI Agent
10 iterations of an AI agent's system prompt: why the model copies examples, why a condition works better than an adjective, and how to test a prompt so the result doesn't lie.
Knowledge
We write about what actually works in the enterprise - Temporal, Camunda, UiPath, Agentic AI and process automation, no fluff.
10 iterations of an AI agent's system prompt: why the model copies examples, why a condition works better than an adjective, and how to test a prompt so the result doesn't lie.
8 posts
Business process automation (BPA) lifts team efficiency by 20-50% while cutting errors, costs and turnaround times. How RPA works, which processes are worth handing to bots, and what it costs.
An autonomous system can handle an order, a complaint or a service request in three seconds. Without stable, multichannel customer communication, all that speed stalls on the last metre — on a busy line and across scattered inboxes.
We had one machine, dozens of models, and last week's test suite. What to look at when picking a local model, what to change in Ollama so it stops hanging, and how it performs after a week of use.
We finally wanted to measure how our agent behaves, and check it across four models. Two identical runs produced the same 2/18 score. Five scenarios behaved differently. With the same agent, swapping the judge alone moved the score from 61% to 11%.
A language model doesn't check facts, it just predicts the next word — it can confidently confirm a promotion that ended a month ago. RAG feeds it up-to-date knowledge from documents before it answers, but implementing it is more than plugging files into a database.
The tool call returned the correct conversation state, and the model still wrote that the goal was agreed. How we built explicit dialogue state via tool calling in Spring AI — and two different ways it can fail.
104 lines of Java, Spring AI 2.0 and our own LLM instead of the cloud. What came out of three days, how much of it is code, and how much is the agent's system prompt.
Mass operations: Kafka, a queue, or Temporal? (an ADR)
No results for the selected filters.