Conversation
Some of the best models for long coding sessions are text-only — you pick them for the context window, then hit a wall the moment an image lands in the conversation and the model throws "not multimodal" and wedges the turn. The vision bridge fixes that when the gateway also serves a multimodal model. Set a vision model (saved or via SETFREE_VISION_MODEL) and SetFree starts a local proxy the CLI talks to instead of the gateway directly. The proxy forwards everything unchanged except image content blocks: each is captioned by the vision model and replaced with text before the request reaches the main model. The text model never sees a raw image block, so it never errors. Captions are cached to disk (NDJSON, keyed by image bytes) so an image is described once ever, not once per turn. This is the one deliberate exception to SetFree's "step aside, never sit in the request path" rule. It's strictly opt-in — no vision model means no proxy and a launch byte-identical to today. The proxy is localhost, lives only as long as the CLI does, and never touches auth or licensing. The CLI binary is still the one you installed, unmodified. Ported the captioning approach from a forked-CLI implementation (claude-local's visionBridge.ts): Anthropic image-block detection at the top level and nested one level inside tool_result content (the easy-to-miss case that reproduces the not-multimodal error), the captioning call forcing chat_template_kwargs.enable_thinking=false (a vLLM passthrough that keeps Qwen3.5 from burning its token budget on thinking and returning empty), and bounded-concurrent prewarm before the serial rewrite (each caption is an ~8s round-trip; serial it looks like a hang). Lifecycle: the vision path runs the CLI as a child (launcher.Run, new — signal-forwarded) instead of exec-replacing (launcher.Launch), so the parent survives to tear the proxy down when the CLI exits. An idle-timeout (default 1h) is only a backstop for a proxy orphaned by an abnormal parent death; it never culls a live session that's between turns. The non-vision path is unchanged. Covers the Anthropic /v1/messages path (claude and vscode adapters). Codex uses a different wire format (OpenAI Responses API) and will follow. Config: [vision] model = "qwen3.5" # presence enables the bridge # base_url and api_key optional, default to the main gateway Env overrides: SETFREE_VISION_MODEL, SETFREE_VISION_BASE_URL, SETFREE_VISION_API_KEY, SETFREE_VISION_OFF, SETFREE_VISION_CONCURRENCY, SETFREE_VISION_IDLE_TIMEOUT. The vision API key lives in the secrets store under "vision", never in config.toml. Limitation vs a forked-CLI bridge: the proxy can't keep the original image for an on-demand "look closer" re-query, so the text model works from the caption. That's the trade for staying out of the binary.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reopens #1, rebased on current
main. Same code you approved there — the only change is theapp.goswitch, which now carries both the newusagecase and the vision one. Green on all five jobs, macOS included, now that #4 is in: https://github.com/ihubanov/setfree/actions/runs/35746015027The problem
Some of the best models for long coding sessions are text-only. You pick one for the context window, then an image lands in the conversation and the turn dies with a "not multimodal" error. I hit this daily running Claude Code on a text-only model with a 1M window, while the same gateway also serves a multimodal model that could have described the image.
How it works
Set a vision model and SetFree starts a small proxy on localhost, then points the CLI at it instead of the gateway. The proxy forwards everything unchanged except image blocks in
POST /v1/messages: each image goes to the vision model, comes back as text, and the main model receives a caption instead of an image it can't read. Captions are cached on disk, so an image is described once ever, not once per turn.or
SETFREE_VISION_MODEL=qwen3.5 setfree claudefor a single run.Covers the Anthropic
/v1/messagespath, soclaudeandvscode. Codex uses the OpenAI Responses API with a different image shape; I left it out to keep this reviewable and will follow up separately.On the philosophy
This is the one place SetFree sits in the request path, so it's strictly opt-in: with no vision model set, no proxy starts and the launch is byte-identical to today. The proxy is the same setfree binary re-invoked as a subcommand, it binds to
127.0.0.1on an ephemeral port, and it exits with the CLI. The CLI binary is still the one you installed, unmodified, and nothing touches auth or licensing.Notes for review
tool_resultcontent, warming captions concurrently before the rewrite (serial captioning of N images looks like a hang), and an append-only NDJSON cache so parallel writers don't clobber each other.chat_template_kwargs.enable_thinking=false. Qwen3.5 is a reasoning model and otherwise spends its budget in<think>and returns nothing. It's a vLLM passthrough that other endpoints ignore.[Attached image — vision description unavailable.]rather than failing the request, so a flaky vision endpoint never breaks a turn.tool_resultcase, the rewrite, cache hits, and proxy passthrough.Happy to split this up or trim the scope if it's easier to review in pieces.