Vision
Give fx a screenshot, diagram, or photo when the prompt needs visual context.
Attach an image
In fx, run /image <path> to add an image to the prompt. You can also drag a local image into the terminal; fx attaches the path it pastes. On macOS, /paste adds an image from the clipboard. /images shows the images you have attached, and /images clear removes them.
For a one-off request, pass the image with --image:
fx ask --image <path> "Describe this interface."Repeat --image to include more than one image. Editors using ACP can send images with session/prompt; see ACP server.
Supported inputs
Prompt attachments support PNG, JPEG, GIF, and WebP files up to 3.75 MiB each. On macOS, fx also tries to resize larger files up to 20 MiB before upload. See ACP server for inline editor image limits.
The agent can also read images up to 3.75 MiB with read_file. Models with image input receive the pixels; other models receive guidance to use the vision tool when available.
How vision works
fx uses the selected model when it can accept image input. The image travels with the prompt and the model reads it directly.
With AI Gateway, fx has another route for models that cannot accept images. It sends the image to a fixed helper model, google/gemini-2.5-flash, and gives the selected model the resulting visual evidence. This is a separate model request with its own usage and cost; see Additional model requests.
ChatGPT and Grok subscriptions use native image input. For a custom connection, set supports_vision: true for a model that accepts images.
Trust
fx treats instructions found inside an image as untrusted content.