Google/gemini-3-flash-preview/Released Dec 2025

Gemini 3 Flash

Low-latency multimodal model, 1M context

Text

About Gemini 3 Flash

gemini-3-flash-preview is a Multimodal LLM. It excels in agentic workflows, multi-turn chat, coding assistance, and interactive tasks due to its lower latency and near-Pro reasoning compared to larger Gemini variants.

Some other noteworthy features of gemini-3-flash-preview include configurable thinking levels (minimal, low, medium, high), structured output, tool use, automatic context caching, and support for multimodal inputs like text, images, audio, video, and PDFs.

MetricValue
Context Length1M tokens
MultilingualYes

Ready to build with Gemini 3 Flash?

Try Gemini 3 Flash in the Workbench to prompt it, compare outputs, and iterate on prompts without writing any code. When you're ready to ship, call the same model from our API and build your own apps on top of it.

curl -sSf -X POST https://hub.oxen.ai/api/ai/chat/completions \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $OXEN_API_KEY" \
    -d '{
  "model": "gemini-3-flash-preview",
  "messages": [
    {
      "role": "user",
      "content": "Try sending a message."
    }
  ]
}'

API endpoint

See the API reference for request and response formats.

POSThttps://hub.oxen.ai/api/ai/chat/completions