About Gemini 3.8 Flash
gemini-3.8-flash is a Multimodal LLM. It is Google's Flash-tier model for long-horizon software engineering, autonomous agents, and multi-step enterprise workflows, with gains over Gemini 3.7 Flash across software engineering, agentic tasks, and specialized-domain reasoning. Google reports 54.9% on HLE-Verified and says it outperforms most larger models on DeepSWE v1.1, its long-horizon engineering benchmark.
Some other noteworthy features of gemini-3.8-flash include configurable thinking levels (low, medium, high; minimal is not supported and returns an error), structured output, function calling, code execution, context caching, search grounding, URL context, file search, computer use (preview), and batch processing. It accepts text, images, audio, video, and PDFs and outputs text only. On hard tasks it tends to run extra reasoning steps and iterative tool calls, so token usage per task can be higher than 3.7 Flash, especially at higher thinking levels. The Live API, image generation, and audio generation are not supported.
| Metric | Value |
|---|---|
| Context Length | 1M tokens |
| Max Output | 64K tokens |
| Multilingual | Yes |
Ready to build with Gemini 3.8 Flash?
Try Gemini 3.8 Flash in the Workbench to prompt it, compare outputs, and iterate on prompts without writing any code. When you're ready to ship, call the same model from our API and build your own apps on top of it.
curl -sSf -X POST https://hub.oxen.ai/api/ai/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OXEN_API_KEY" \
-d '{
"model": "gemini-3-8-flash",
"messages": [
{
"role": "user",
"content": "Try sending a message."
}
]
}'