A delegation waits on its own sub-thread
What went wrong
The function invoker runs the tool calls of one model turn concurrently
(AllowConcurrentInvocation = true in ChatClientAgentFactory). A delegate_to_agent call waits in
two steps:
- It waits for the chat's delegation stream (
IAgentChat.Delegations) to report that the sub-thread was created (Dispatched) or could not be created (Failed). - It then reads that sub-thread's node until the round ends, bounded by
WaitForDelegationResult's 10-minute backstop.
Step 1 took the next Dispatched or Failed on the stream and did not check whose it was. When two
delegations ran in the same turn, both calls subscribed before either sub-thread existed. The first
sub-thread to be created then settled both calls. The effects:
- The second call answered with the first call's summary.
- The second call's own sub-thread ran with nobody waiting on its result.
- If one create failed, the other call was reported as having failed to start.
Nothing errored and nothing was logged. The model simply got a wrong tool result back. A related
problem: the stream was a bare Subject, and concurrent delegations call OnNext from different
threads, which breaks the ordering every subscriber's Where and Take relies on.
The fix
delegate_to_agentmints a call id (DelegationTool.NewCallId, a full GUID so that two calls in flight cannot share one) and passes it to the delegation it starts (ExecuteDelegationAsync, through theexecuteAsyncdelegate).- Every lifecycle event for that call (
Dispatched,Failed,Terminal) carries the id. - The wait resolves only on its own call's
DispatchedorFailed(DelegationTool.LifecycleOf). AgentChatClient's delegation stream isSubject.Synchronize(...), so concurrent emits are serialised.
The CreateDelegationTools overload whose executeAsync takes no call id still exists, marked
[Obsolete]. It keeps the old uncorrelated wait, because an implementation that never receives the id
cannot stamp it. That wait is only right while one call is waiting for its start, and [Obsolete]
warns only a caller that is recompiled. So a tool built by that overload refuses a delegate_to_agent
call that arrives while another one is still waiting for its Dispatched or Failed: the call
returns an Error: … tool result naming the overlap, and the model can call again. Once the first
call's start has settled, the next call is accepted.
ConcurrentDelegationCorrelationTest covers this. It runs two calls on one chat and reports the second
call's failure first. With the call-id filter removed, call A reports B's error and the test fails. A
foreign event, emitted first, is the negative control: it settles neither call.
TheUncorrelatedOverload_RefusesAnOverlappingStart covers the retired overload: the overlapping call
is refused, the first call still reports its own result, and a call made after that start settled is
accepted.
What this does not establish
MeshWeaver#5915's own incident was a delegate_to_agent call that ran 23.6 minutes until the
30-minute round cap cut it. By the code, a dispatched delegation returns within the 10-minute backstop,
and a create that never answers fails within the request timeout. That leaves the following open:
- Which phase that call was really waiting in. The round's
[Delegation:…]lines would show it, and they were outside the Loki window that could be read. - Whether the crossing above played a part. It can hand a call the wrong sub-thread, but by itself it cannot make a call wait past the backstop.
Nothing here extends the round cap or adds a bound.