← News
Jérémie Doucy·

Why our agents classify each turn after they answer

Every time an agent on the Bayes Platform replies, the platform needs a few facts about that reply. It needs a title for the conversation, the categories our partners track in their analytics dashboards, and the knowledge base passages the reply relies on, which feed the sources panel next to the answer.

For two months we computed these inside the answering generation itself, through a tool the model had to call. In mid-September we moved them to a dedicated call that runs once the reply is complete. This post explains both designs and what made us switch.

A mandatory tool in the loop

Our agents already call tools: they search the knowledge base, fill forms, hand the conversation to sub-agents. Adding one more felt natural. It took the title, the categories and the ids of the cited passages as arguments, and when the model called it alongside its answer, classification cost nothing.

On a production-sized prompt of about 15,000 tokens, though, smaller models rarely volunteered the call. At temperature 0 the rate ranged from 0/8 on the Gemini 3.5 family to 3/8 on Gemini 3.6 Flash. On a simple greeting there is nothing to summarize, and the model reasonably judged the tool irrelevant.

Forcing the call does not help. With toolChoice: "required", every provider we tested (Gemini, Gemma 4 on vLLM, Claude) produced the call and dropped the text answer. There is no single-generation mode for “call this tool and also answer”.

So we built a hybrid. The tool was renamed mandatory_tool, and that rename alone took the call rate on Gemini 3.6 Flash from 3/8 to 8/8, the largest effect of the whole prompt campaign. When the model still skipped it, one forced catch-up generation ran after the stream. The user already had the answer by then, so losing the text of that extra generation did not matter.

It worked. Every turn got its report, on every provider.

What we got wrong

We had asked the answering model to do a job unrelated to the conversation, and smaller models made the cost visible quickly.

The clearest symptom was leaked calls. A model would sometimes write the call out as text instead of emitting it, and the user would see mandatory_tool{…} in the middle of a reply. Every pattern our output sanitizer learned to strip between July and September was a leak of that same tool, and each fix was one more regex.

The tool also crowded the turn. With several obligatory calls in one turn (a form to fill, a handoff signal, the report), models started asking two questions at once or skipping a form field. On the code side, half of the streaming pipeline existed to keep one report alive. It needed dynamic schema getters and execution accounting, a special stop condition, and a protocol section appended to every agent’s prompt.

Answer first, classify after

The new design follows one rule: the answering loop only carries tools that interact with the conversation. Anything the platform wants to know about the reply is computed afterwards, by two mechanisms.

The two phases of a turnDuring the replyAfter the replyUser messageAnswering loopknowledge base, forms, MCP, sub-agentsReply streamed and storedcitation markers removedClassification callstructured output, temperature 0SourcesConversation titleCategoriesHandoff conclusion[c1]fallback, whennothing is cited

Sources as inline citations

When the knowledge base search returns passages, it gives each one a short alias (c1, c2…). The model cites a passage by writing its alias in brackets right after the sentence it supports: [c1], [c1, c3]. It copies two characters it has just read. This is the same grounding pattern Perplexity, Gemini grounding, Claude and Cohere citations use.

A streaming extractor removes the markers from the text the user sees and from the stored message, and records which aliases were cited. The tricky part is that a marker can arrive split across two stream deltas, so the extractor holds back everything from an unclosed [ until the ] arrives or the stream ends:

const openBracket = pending.lastIndexOf("[")
if (
  openBracket !== -1 &&
  pending.length - openBracket < MAX_MARKER_LENGTH &&
  !pending.includes("]", openBracket)
) {
  safeUntil = openBracket
}

An open bracket more than 48 characters back is treated as prose and released. The pattern also tolerates the variants smaller models produce, such as extra spaces, ; separators or upper case.

One structured-output call

Once the reply is complete and stored, a single structured-output call produces the title and the categories. It also attributes sources, but only when passages were retrieved and the model cited none of them inline.

The classifier never sees the agent’s system prompt, which was exactly what made the tool call unreliable. It reads a short prompt that is the same for every agent. The prompt holds the latest exchange and the six previous messages (truncated to 600 characters), then the category list and passage excerpts when needed.

We use structured output rather than tool calling here because there is no longer any text to preserve, so the constraint that stopped us from forcing a tool call disappears. It also gives us a stronger guarantee on Gemma 4. When we serve it with vLLM, a JSON schema is compiled into a grammar that masks every token outside the schema, so the model cannot produce a category that is not in the list. Tool calls get no such constraint: vLLM lets Gemma write the call in its own format and parses the text once it is generated, so an invalid value only shows up after the fact.

The schema is built for each turn and only declares the fields that turn needs. Closed lists become enums, so a constrained decoder cannot invent a category or a passage id:

if (availableCategoryNames.length > 0) {
  properties.categoryNames = {
    type: "array",
    items: { type: "string", enum: availableCategoryNames },
    maxItems: MAX_CATEGORIES,
    description:
      "The complete set of categories that apply to the conversation as a whole, including earlier turns. Empty when none applies.",
  }
}

The call runs at temperature 0 on the agent’s own model, and it is best-effort. The user already has the reply, so a failure is logged loudly and never breaks the stream.

Categories deserve a note because they drive the analytics. We ask for the complete set that applies to the whole conversation, at most five, and replace the session’s categories with it on every turn. A turn that gets the categories slightly wrong is corrected by the next one instead of piling up. The results are still logged under the old tool names, so stored conversations, the activity timeline and the dashboards did not change.

A step after every turn

Taking classification out of the loop gave us a guaranteed step after every turn, where we can ask new questions without touching how the reply is generated. We used it within a week.

When a sub-agent takes over a conversation, for example to walk the user through a questionnaire, the schema gains two fields: taskConcluded says whether the reply ends the sub-agent’s part, and handoffSummary gives the parent agent two to four sentences on what was collected. If the sub-agent forgets to call its conclusion tool, the platform applies the conclusion from the classifier’s reading. In the old design, that would have been one more obligation inside a loop that already carried too many.

Results

The new design costs one small extra generation per turn, including on models that used to produce the report for free. Models that never volunteered it were already paying for the forced catch-up.

We ran the live regression suites on the full pipeline on September 16:

Suite Result
RAG turn, 7 models 7/7: grounded answer, no marker left in the text, sources and metadata logged once
Large prompt, 7 models × 4 turns 28/28: one classification per turn, categories from the list, a title on every turn with content

All seven models (Gemma 4 and six Gemini variants from 2.5 Flash to 3.8 Flash) cited their sources inline on the RAG scenario, so the classifier’s fallback attribution was never needed there. On the large prompt, every model returned no title for the greeting and the expected category on the other turns.

The pipeline also lost its dynamic getters, its execution accounting and its prompt epilogue, and the output sanitizer went back to protecting the tools that talk to the user.

What’s next

The citation markers carry the exact position of each source in the reply. Version one strips them, and the next step is to show them as footnotes. A safety signal is a better fit as a field of the classification than as a tool the model might forget, and that is the next field we plan to add. Finally, the classifier runs on each agent’s own model. A smaller dedicated model would be cheaper, but it has to be weighed partner by partner against their data residency requirements, so for now the agent’s model stays the default.