Skip to content

Add a vision bridge for text-only main models - #6

Open
ihubanov wants to merge 1 commit into
mindsdb:mainfrom
ihubanov:feat/vision-bridge-v2
Open

ihubanov wants to merge 1 commit into
mindsdb:mainfrom
ihubanov:feat/vision-bridge-v2

Conversation

@ihubanov

Copy link
Copy Markdown
Contributor

Reopens #1, rebased on current main. Same code you approved there — the only change is the app.go switch, which now carries both the new usage case and the vision one. Green on all five jobs, macOS included, now that #4 is in: https://github.com/ihubanov/setfree/actions/runs/35746015027

The problem

Some of the best models for long coding sessions are text-only. You pick one for the context window, then an image lands in the conversation and the turn dies with a "not multimodal" error. I hit this daily running Claude Code on a text-only model with a 1M window, while the same gateway also serves a multimodal model that could have described the image.

How it works

Set a vision model and SetFree starts a small proxy on localhost, then points the CLI at it instead of the gateway. The proxy forwards everything unchanged except image blocks in POST /v1/messages: each image goes to the vision model, comes back as text, and the main model receives a caption instead of an image it can't read. Captions are cached on disk, so an image is described once ever, not once per turn.

[vision]
model = "qwen3.5"
# base_url and api_key are optional; they default to your main gateway

or SETFREE_VISION_MODEL=qwen3.5 setfree claude for a single run.

Covers the Anthropic /v1/messages path, so claude and vscode. Codex uses the OpenAI Responses API with a different image shape; I left it out to keep this reviewable and will follow up separately.

On the philosophy

This is the one place SetFree sits in the request path, so it's strictly opt-in: with no vision model set, no proxy starts and the launch is byte-identical to today. The proxy is the same setfree binary re-invoked as a subcommand, it binds to 127.0.0.1 on an ephemeral port, and it exits with the CLI. The CLI binary is still the one you installed, unmodified, and nothing touches auth or licensing.

Notes for review

  • Captioning is ported from a working implementation I've been running for months, including the details that matter: recursing one level into tool_result content, warming captions concurrently before the rewrite (serial captioning of N images looks like a hang), and an append-only NDJSON cache so parallel writers don't clobber each other.
  • The vision call sends chat_template_kwargs.enable_thinking=false. Qwen3.5 is a reasoning model and otherwise spends its budget in <think> and returns nothing. It's a vLLM passthrough that other endpoints ignore.
  • A failed caption degrades to [Attached image — vision description unavailable.] rather than failing the request, so a flaky vision endpoint never breaks a turn.
  • Two things a proxy can't do that an in-process hook could: keep the original image for an on-demand "look closer" re-query, and let the text model ask a follow-up about a detail the caption missed. That's the price of staying out of the binary.
  • Tests cover config precedence, block detection including the nested tool_result case, the rewrite, cache hits, and proxy passthrough.

Happy to split this up or trim the scope if it's easier to review in pieces.

Some of the best models for long coding sessions are text-only — you pick
them for the context window, then hit a wall the moment an image lands in
the conversation and the model throws "not multimodal" and wedges the turn.

The vision bridge fixes that when the gateway also serves a multimodal
model. Set a vision model (saved or via SETFREE_VISION_MODEL) and SetFree
starts a local proxy the CLI talks to instead of the gateway directly.
The proxy forwards everything unchanged except image content blocks:
each is captioned by the vision model and replaced with text before the
request reaches the main model. The text model never sees a raw image
block, so it never errors. Captions are cached to disk (NDJSON, keyed by
image bytes) so an image is described once ever, not once per turn.

This is the one deliberate exception to SetFree's "step aside, never sit
in the request path" rule. It's strictly opt-in — no vision model means
no proxy and a launch byte-identical to today. The proxy is localhost,
lives only as long as the CLI does, and never touches auth or licensing.
The CLI binary is still the one you installed, unmodified.

Ported the captioning approach from a forked-CLI implementation
(claude-local's visionBridge.ts): Anthropic image-block detection at the
top level and nested one level inside tool_result content (the easy-to-miss
case that reproduces the not-multimodal error), the captioning call forcing
chat_template_kwargs.enable_thinking=false (a vLLM passthrough that keeps
Qwen3.5 from burning its token budget on thinking and returning empty),
and bounded-concurrent prewarm before the serial rewrite (each caption is
an ~8s round-trip; serial it looks like a hang).

Lifecycle: the vision path runs the CLI as a child (launcher.Run, new —
signal-forwarded) instead of exec-replacing (launcher.Launch), so the
parent survives to tear the proxy down when the CLI exits. An idle-timeout
(default 1h) is only a backstop for a proxy orphaned by an abnormal parent
death; it never culls a live session that's between turns. The non-vision
path is unchanged.

Covers the Anthropic /v1/messages path (claude and vscode adapters). Codex
uses a different wire format (OpenAI Responses API) and will follow.

Config:
  [vision]
  model = "qwen3.5"        # presence enables the bridge
  # base_url and api_key optional, default to the main gateway

Env overrides: SETFREE_VISION_MODEL, SETFREE_VISION_BASE_URL,
SETFREE_VISION_API_KEY, SETFREE_VISION_OFF, SETFREE_VISION_CONCURRENCY,
SETFREE_VISION_IDLE_TIMEOUT. The vision API key lives in the secrets store
under "vision", never in config.toml.

Limitation vs a forked-CLI bridge: the proxy can't keep the original image
for an on-demand "look closer" re-query, so the text model works from the
caption. That's the trade for staying out of the binary.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant