Repository navigation
Replies: 1 comment
|
The key point is that the native tool iteration loop is not owned by
Because of that, I would not use
If you need visibility or context management between model calls inside a long tool run, the most useful interception point is the A simplified pipeline looks conceptually like this: If you place a custom For example: public sealed class ContextManagingChatClient : DelegatingChatClient
{
public ContextManagingChatClient(IChatClient innerClient)
: base(innerClient)
{
}
public override async Task<ChatResponse> GetResponseAsync(
IEnumerable<ChatMessage> messages,
ChatOptions? options = null,
CancellationToken cancellationToken = default)
{
var managedMessages = CompressIfNeeded(messages);
return await base.GetResponseAsync(
managedMessages,
options,
cancellationToken);
}
private static IEnumerable<ChatMessage> CompressIfNeeded(
IEnumerable<ChatMessage> messages)
{
// token counting / summarization / pruning strategy
return messages;
}
}Then compose the pipeline so that the context-management client wraps the actual provider client before the function-invocation layer performs its repeated calls. Conceptually: That ordering matters. If your middleware wraps the you normally see the outer request, but not every internal model round performed by the function-calling loop. So regarding your questions:
For long-running agents I would therefore treat context management as middleware: This also keeps token-budget enforcement independent from the agent implementation itself. One additional safeguard worth using is So my recommended approach would be:
If the framework eventually adds a native before/after-tool-round callback, that would be cleaner for observability, but the |
Uh oh!
There was an error while loading. Please reload this page.
Hi - I have a big issue trying to manage context during long runs where the gframework internally manages all tool iterations.
Hoping someone can point me in the right direction
an event on ChatClientAgent?
call) makes it useless for long tool runs.
I added post-stream compression, but it's not enough, and I cannot see a way to implement a hook that will help prevent a single 15-tool run from killing the conversation.
Thanks.
All reactions