Record a voice memo, drop in an image, or just type. Audio goes straight into the model — no transcription step in front, no separate vision endpoint beside it. Served by Plubo on an OpenAI-compatible API. No key needed here.
Audio is capped at 30 seconds — that is the model's limit, not the demo's. Gemma 4's audio ability is speech recognition and speech translation, per Google's model card; it is not a tone or emotion classifier.
| Model | Context | Input | $/1M in · out | Image · audio $/1M |
|---|---|---|---|---|
gemma-4-12b-it | 262,144 | text · image · audio · video | — | — |
gemma-4-e4b-it | 131,072 | text · image · audio · video | — | — |
gemma-4-e2b-it | 131,072 | text · image · audio · video | — | — |