Skip to content

feat(workflow-eval): add batch tasks and scorers - #879

Merged
Abhijeet Prasad (AbhiPrasad) merged 4 commits into
mainfrom
abhi/implement-python-sdk-a3bb3ad9
Oct 5, 2026
Merged

Abhijeet Prasad (AbhiPrasad) merged 4 commits into
mainfrom
abhi/implement-python-sdk-a3bb3ad9

Conversation

@AbhiPrasad

Copy link
Copy Markdown
Member

Adds WorkflowBatchTask and WorkflowBatchScorer to the Python workflow eval API, matching the batch workflow support in the JavaScript SDK.

Batch processors submit up to batching.max_size cases at once. Each submitted WorkflowBatchItem has a provider-safe custom_id and an .item containing the original task or scorer item. Collectors return WorkflowBatchItemResult values keyed by custom_id; a reported error or missing result fails only that item. Failed tasks skip their scorers, while scorer failures are recorded on the scorer span.

Example

from braintrust.workflow_eval import (
    WorkflowBatchItemResult,
    WorkflowBatchTask,
    WorkflowBatchingOptions,
    WorkflowSubmissionCompletionWebhook,
    WorkflowTaskResult,
    define_workflow_eval,
)

async def submit(items, context):
    requests = [
        {"custom_id": item.custom_id, "input": item.item.input}
        for item in items
    ]
    return await provider.create_batch(requests)

async def collect(batch, context):
    async for row in provider.results(batch["id"]):
        yield WorkflowBatchItemResult(
            custom_id=row["custom_id"],
            result=WorkflowTaskResult(output=row["output"]),
        )

task = WorkflowBatchTask(
    batching=WorkflowBatchingOptions(max_size=10_000),
    submit=submit,
    completion=WorkflowSubmissionCompletionWebhook(lambda batch, _ctx: batch["id"]),
    collect=collect,
)

evaluation = define_workflow_eval(
    "Support QA", store=store, data=data, task=task, scores=[batch_scorer]
)

Batch scorers use the same callbacks and support WorkflowBatchingOptions(max_size=..., max_wait_ms=...). A partial scorer batch is submitted after its wait window or when no more task results can arrive. The wait is checked as the eval advances; it does not start a background timer.

flowchart LR
    A[Eval cases] --> B[Batch task submissions]
    B --> C[Per-item task results]
    C --> D[Batch scorer submissions]
    D --> E[Per-item scores and eval rows]
    F[Poll or webhook] --> B
    F --> D
Loading

A webhook can resume a run with the provider batch ID alone:

await evaluation.process_submission_result(external_id=provider_batch_id)

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-05T23:25:04.535258Z 49e76af New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ce536fce61

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".

Comment thread py/src/braintrust/workflow_eval.py Outdated
Comment thread py/src/braintrust/workflow_eval.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 8ec2da9d31

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".

Comment thread py/src/braintrust/workflow_eval.py
Comment thread py/src/braintrust/workflow_eval.py Outdated
Comment thread py/src/braintrust/workflow_eval.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d81dea7f68

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".

Comment thread py/src/braintrust/workflow_eval.py Outdated
Comment thread py/src/braintrust/workflow_eval.py Outdated
Comment thread py/src/braintrust/workflow_eval.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 49e76af819

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "Codex (@codex) review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "Codex (@codex) address that feedback".


DEFAULT_MAX_CONCURRENCY = 10
DEFAULT_REDIS_TTL_MS = 1_000 * 60 * 60 * 24 * 7
_BATCH_LEASE_TTL_MS = 5 * 60 * 1000

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep the retry lease alive through provider submission

The fixed five-minute expiration is fresh evidence beyond the earlier concurrent-retry issue: when a user-defined submit callback remains in flight longer than _BATCH_LEASE_TTL_MS, another poll can acquire the expired lease while the persisted placeholder is still failed and invoke processor.submit for the same items again. Both callbacks may create provider jobs and race to persist their submission data, so renew the lease while provider work is running or atomically transition the record out of failed before invoking submit.

Useful? React with 👍 / 👎.

Comment on lines 1400 to 1401
if submission.get("external_id") is not None:
await self._write(run_id, "external", submission["id"], submission["external_id"])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Repair the external locator when replaying progress

If the worker crashes or the store transiently fails after the submission record is written with its external ID but before _record_external_locator succeeds, this replay path only restores the per-run external key and progress counters. An external-ID-only webhook will therefore continue to raise No submission matches this external_id, and webhook submissions are never polled, leaving the run stuck unless the caller retained the run ID; replay the global locator from the persisted submission as well.

Useful? React with 👍 / 👎.

Comment on lines +1509 to +1513
await self._write(
run_id,
"task-result",
{"output": None, "metadata": {}, "tags": None, "error": error},
item_id,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve case metadata on failed batch task items

When a batch task entry reports an error or is missing, this synthesized task result replaces the case metadata and tags with empty values. _summary later reads these fields from task-result, so no_send_logs=True evaluations lose the dataset metadata and tags specifically on failed rows, unlike the ordinary evaluator and successful batch rows; retain the case metadata and tags while adding the error.

Useful? React with 👍 / 👎.

@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) merged commit c539c57 into main Oct 5, 2026
83 checks passed
@AbhiPrasad
Abhijeet Prasad (AbhiPrasad) deleted the abhi/implement-python-sdk-a3bb3ad9 branch October 5, 2026 23:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants