RUN THIS WORKFLOW NOW ON FLOYO!
ABOUT THE WORKFLOW
Read an Image and Answer
Upload an image and type a question or instruction. The model reads what is in the image and answers in plain language. You can also add an audio clip to transcribe or describe alongside the image.
Model
Gemma 4 E4B by Google DeepMind. An open-weights multimodal model built from Gemini 3 research that takes text, image, and audio and writes a text response. Strong at description, analysis, transcription, and question answering, with a configurable thinking mode.
Description
concept
model
ai
image
comfyui
workflow
gemma
i2t
opensource
floyo
google deepmind
multimodal model
ask about image
image reading
gemma 4 e4b
Details
Downloads
66
Platform
CivitAI
Platform Status
Available
Created
6/25/2026
Updated
8/11/2026
Deleted
-
