Repository navigation
feat(workflow-eval): add batch tasks and scorers - #879
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ce536fce61
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "Codex (@codex) review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 8ec2da9d31
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "Codex (@codex) review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d81dea7f68
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "Codex (@codex) review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 49e76af819
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "Codex (@codex) review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".
|
|
||
| DEFAULT_MAX_CONCURRENCY = 10 | ||
| DEFAULT_REDIS_TTL_MS = 1_000 * 60 * 60 * 24 * 7 | ||
| _BATCH_LEASE_TTL_MS = 5 * 60 * 1000 |
There was a problem hiding this comment.
Keep the retry lease alive through provider submission
The fixed five-minute expiration is fresh evidence beyond the earlier concurrent-retry issue: when a user-defined submit callback remains in flight longer than _BATCH_LEASE_TTL_MS, another poll can acquire the expired lease while the persisted placeholder is still failed and invoke processor.submit for the same items again. Both callbacks may create provider jobs and race to persist their submission data, so renew the lease while provider work is running or atomically transition the record out of failed before invoking submit.
Useful? React with 👍 / 👎.
| if submission.get("external_id") is not None: | ||
| await self._write(run_id, "external", submission["id"], submission["external_id"]) |
There was a problem hiding this comment.
Repair the external locator when replaying progress
If the worker crashes or the store transiently fails after the submission record is written with its external ID but before _record_external_locator succeeds, this replay path only restores the per-run external key and progress counters. An external-ID-only webhook will therefore continue to raise No submission matches this external_id, and webhook submissions are never polled, leaving the run stuck unless the caller retained the run ID; replay the global locator from the persisted submission as well.
Useful? React with 👍 / 👎.
| await self._write( | ||
| run_id, | ||
| "task-result", | ||
| {"output": None, "metadata": {}, "tags": None, "error": error}, | ||
| item_id, |
There was a problem hiding this comment.
Preserve case metadata on failed batch task items
When a batch task entry reports an error or is missing, this synthesized task result replaces the case metadata and tags with empty values. _summary later reads these fields from task-result, so no_send_logs=True evaluations lose the dataset metadata and tags specifically on failed rows, unlike the ordinary evaluator and successful batch rows; retain the case metadata and tags while adding the error.
Useful? React with 👍 / 👎.
Adds
WorkflowBatchTaskandWorkflowBatchScorerto the Python workflow eval API, matching the batch workflow support in the JavaScript SDK.Batch processors submit up to
batching.max_sizecases at once. Each submittedWorkflowBatchItemhas a provider-safecustom_idand an.itemcontaining the original task or scorer item. Collectors returnWorkflowBatchItemResultvalues keyed bycustom_id; a reported error or missing result fails only that item. Failed tasks skip their scorers, while scorer failures are recorded on the scorer span.Example
Batch scorers use the same callbacks and support
WorkflowBatchingOptions(max_size=..., max_wait_ms=...). A partial scorer batch is submitted after its wait window or when no more task results can arrive. The wait is checked as the eval advances; it does not start a background timer.flowchart LR A[Eval cases] --> B[Batch task submissions] B --> C[Per-item task results] C --> D[Batch scorer submissions] D --> E[Per-item scores and eval rows] F[Poll or webhook] --> B F --> DA webhook can resume a run with the provider batch ID alone: