Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Qwen3.8-Omni-Flash vs Gemini 3.8 Flash

This is the comparison Alibaba picked itself. Qwen3.8-Omni-Flash, released September 18, 2026, claims audio-visual performance close to Gemini 3.8 Flash and overall audio performance above it, so the question is whether a five times cheaper rate card buys you the same job. Qwen bills $0.15 per million input tokens and $0.47 output with cache hits at $0.016. Gemini 3.8 Flash, released September 2, bills $0.75 and $3.75, and Google printed the expiry in a launch-post footnote: both sides double to $1.50 and $7.50 on January 1, 2027. So the gap today is five times on input and eight times on output, and it widens to ten and sixteen after the new year. Both carry roughly 1 million token context windows and both take text, image, audio, and video. The differences are in what comes out and who says so. Qwen returns text only, so speech generation needs a separate model, while Gemini handles more output paths on one endpoint. Qwen has published no independent benchmark at all: the 29-evaluation average, the 36.5 point WildClawBench-MM gain, and the OmniVideoBench lift from 63.4 to 67.8 are all vendor-reported with Qwen's own estimation method, whereas Artificial Analysis publishes an index and a cost-per-task figure for Gemini 3.8 Flash. Qwen's genuine architectural claim is agentic perception: it decides which parts of a long video to watch rather than reading it front to back, which it says cuts token use about 45.7 percent while raising accuracy. If that holds on your footage, the effective gap is wider than the per-token gap. There is one more axis that is not about quality. Gemini runs on Vertex and AI Studio with the compliance surface that implies; Qwen3.8-Omni-Flash runs on Alibaba Cloud Model Studio in six regions, and Alibaba was one of six labs named in the September 8, 2026 joint CISA, NSA, and FBI advisory on model distillation. That does not change a benchmark, but it will change some procurement reviews.

Head-to-Head Specs

SpecQwen3.8-Omni-FlashGemini 3.8 Flash
ProviderAlibabaGoogle
Input Price$0.15/1M$0.75/1M
Output Price$0.47/1M$3.75/1M
Context Window1M1.0M
Released2026-092026-09
Capabilitiestext, vision, audio, video, tool-use, code, reasoningtext, vision, audio, video, tool-use, code, reasoning

Category Breakdown

Input pricing todayQwen3.8-Omni-Flash

Qwen is $0.15 against $0.75 for Gemini, five times cheaper

Output pricing todayQwen3.8-Omni-Flash

Qwen is $0.47 against $3.75 for Gemini, roughly eight times cheaper

Pricing in 2027Qwen3.8-Omni-Flash

Gemini doubles to $1.50/$7.50 on January 1, 2027; Qwen has announced no increase

Cached inputQwen3.8-Omni-Flash

Qwen lists implicit cache hits at $0.016 per 1M tokens

Output modalitiesGemini 3.8 Flash

Qwen3.8-Omni-Flash returns text only and points at Qwen3.5-Omni for generated speech

Evidence qualityGemini 3.8 Flash

Artificial Analysis publishes an independent index and cost per task for Gemini; every Qwen figure is vendor-reported

Long video efficiencyQwen3.8-Omni-Flash

Agentic perception cuts OmniVideoBench token use about 45.7 percent while raising accuracy, on Qwen's own run

Enterprise tooling and complianceGemini 3.8 Flash

Vertex AI and AI Studio against Alibaba Cloud Model Studio, and Alibaba was named in the September 2026 US distillation advisory

Input modalitiesTieTie

Both take text, image, audio, and video

Context windowTieTie

Both sit at roughly 1 million tokens

Choose Qwen3.8-Omni-Flash when:

  • ▸High-volume audio and video analysis where per-hour cost decides the project
  • ▸Agents that watch or listen and then call tools, at flash pricing
  • ▸Multilingual audio across 113 languages and dialects
  • ▸Budgets that have to survive the January 2027 Gemini price step
View Qwen3.8-Omni-Flash details

Choose Gemini 3.8 Flash when:

  • ▸Workloads that need output modalities beyond text on one endpoint
  • ▸Buyers who require an independently measured score before deploying
  • ▸Vertex AI and AI Studio tooling, quotas, and support
  • ▸Procurement that will not clear a lab named in the US distillation advisory
View Gemini 3.8 Flash details

Frequently Asked Questions

Which is better, Qwen3.8-Omni-Flash or Gemini 3.8 Flash?

It depends on your use case. Qwen3.8-Omni-Flash from Alibaba excels at high-volume audio and video analysis where per-hour cost decides the project, while Gemini 3.8 Flash from Google is better for workloads that need output modalities beyond text on one endpoint. See the full comparison above for detailed benchmarks and pricing.

How much does Qwen3.8-Omni-Flash cost compared to Gemini 3.8 Flash?

Qwen3.8-Omni-Flash costs $0.15 input and $0.47 output per 1M tokens. Gemini 3.8 Flash costs $0.75 input and $3.75 output per 1M tokens.

What is the context window difference between Qwen3.8-Omni-Flash and Gemini 3.8 Flash?

Qwen3.8-Omni-Flash supports 1M tokens, while Gemini 3.8 Flash supports 1.0M tokens.

More Comparisons

Interactive Compare ToolAll ModelsFull Pricing Guide