Gemini 3.8 Flash vs DeepSeek V4 Flash: 59 vs 52, 2.6x Cost

Gemini 3.8 Flash scores 59 on Artificial Analysis vs 52 for DeepSeek V4 Flash, but cost $826 vs $323 to run the same index. Price, speed, and when each wins.

Gemini 3.8 Flash vs DeepSeek V4 Flash: 59 vs 52, 2.6x Cost

Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index at high effort; DeepSeek V4 Flash 0731 scores 52. Running the same index cost $825.83 on Gemini 3.8 Flash and $323.26 on DeepSeek, a 2.6x gap for seven points. At low effort Gemini ties DeepSeek at 52, but AA’s Cost per Task there is $0.24 against $0.11 for DeepSeek.

Google shipped Gemini 3.8 Flash on 2 September 2026 at the same price as 3.7 Flash, and Artificial Analysis called it “the cheapest model at its level of intelligence”. DeepSeek V4 Flash 0731 is the model most people will put next to it, because it is the cheapest per token in the same weight class. Both statements are true. They are answers to different questions, and this post keeps them apart.

TL;DR: Which One Should You Pick?

  • Pick Gemini 3.8 Flash when you need the extra seven points. It scores 59 on the Artificial Analysis Intelligence Index at high effort; DeepSeek’s entire current lineup tops out at 53. It also takes images, audio and files, and streams at roughly 302 output tokens per second.
  • Pick DeepSeek V4 Flash 0731 when a 52 is enough. At that score AA’s Cost per Task is $0.11 for DeepSeek at peak rates against $0.24 for Gemini 3.8 Flash at low effort, and off-peak halves DeepSeek’s again. What you give up is output about 2.2x slower and text-only input.
  • Do not compare sticker prices. DeepSeek wrote 210M output tokens to complete the index; Gemini 3.8 Flash wrote 120M. Per-token price is the smaller half of the bill.
  • Watch 1 January 2027. Gemini 3.8 Flash doubles to $1.50 / $7.50 on that date. Every Gemini number below should be read as a 2026 number.

What Google Changed in Gemini 3.8 Flash

Gemini 3.8 Flash is Google’s third Flash release in six weeks, after 3.6 Flash on 21 July and 3.7 Flash on 13 August. The per-token price did not move: $0.75 per million input tokens and $3.75 output, the same introductory rate as 3.7 Flash, expiring on 31 December 2026.

What did move is how much the model writes. Google’s announcement is unusually direct about it: “3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.” For efficiency-first workloads Google’s advice is to lower the effort level or “continue to rely on Gemini 3.7 Flash, which remains fully supported”.

The same launch introduced Gemini 3.8 Flash Cyber, a vulnerability-detection variant restricted to Google’s Fairwind Program. It is not a general API model and is not compared here.

Quick Specs Comparison

SpecGemini 3.8 FlashDeepSeek V4 Flash 0731
Released2 Sep 202631 Jul 2026
Context window1M tokens1M tokens
Max output64K tokens384K tokens
Input modalitiesText, image, audio, video, fileText only
Reasoning controllow / medium (default) / high; minimal returns an errorthinking on (default) / off; effort low / high / max
WeightsClosedOpen (MIT)
AA Intelligence Index59 high / 57 medium / 52 low52 max effort
AA output speed~302 tokens/s~138 tokens/s
Output tokens to run the AA index120M210M
AA cost to run the index$825.83$323.26

Intelligence Index, token and cost figures are from Artificial Analysis as of 3 September 2026. DeepSeek’s vision-capable build, deepseek-v4-flash-vision-exp, is a separate model and is not the one in this table.

Gemini 3.8 Flash on Artificial Analysis: 59 vs 52, and a Ceiling

Gemini 3.8 Flash leads DeepSeek V4 Flash 0731 by seven points on the Artificial Analysis Intelligence Index, 59 to 52, but only at high effort. Artificial Analysis publishes one composite number per model and reasoning level. Here is the neighbourhood:

ModelAA Intelligence Index
Gemini 3.8 Flash (high)59
Gemini 3.8 Flash (medium, default)57
Gemini 3.7 Flash (high)56
DeepSeek V4 Pro 0813 (max)53
Gemini 3.8 Flash (low)52
DeepSeek V4 Flash 0731 (max)52
Gemini 3.6 Flash (high)52

Two things fall out of that table.

First, the headline gap is seven points, but it is only there at high effort. Ship Gemini 3.8 Flash at its default medium and the gap is five. Turn it down to low and the two models are tied at 52.

Second, the ceiling matters more than the gap. DeepSeek V4 Flash 0731 at 52 is one point below DeepSeek’s own flagship, V4 Pro 0813 at 53. If your task needs a 57 or a 59, there is no DeepSeek model to reach for. The comparison only exists at 52 and below.

Output Tokens: DeepSeek V4 Flash 210M vs Gemini 3.8 Flash 120M

DeepSeek V4 Flash 0731 spent 210M output tokens to complete the Artificial Analysis index; Gemini 3.8 Flash (high) spent 120M, so DeepSeek writes 1.75x as many. Artificial Analysis logs how many output tokens each model spends completing the index. Gemini 3.8 Flash (high) spent 120M, which AA flags as “very verbose” against a median of 71M for comparable models. DeepSeek V4 Flash 0731 spent 210M, “at the higher end compared to other open weight models of similar size (median: 110M)”.

So both are talkative, but DeepSeek writes 1.75x as many, and every one of those tokens is billed at the output rate. Speed compounds it: at 138 tokens per second, DeepSeek needs about 2.2x as long as Gemini 3.8 Flash to emit the same number of tokens, and it emits more of them.

On 3.8 Flash specifically, AA measured a 30% increase in average output tokens per task compared with 3.7 Flash, to 48k, and Cost per Task rising from $0.40 to $0.58 at high effort. Same price list, about 45% higher bill (AA rounds it to roughly 40%). That is the number Google was warning about.

Gemini 3.8 Flash Price vs DeepSeek V4 Flash Price: Sticker, Then Bill

Sticker prices first. Dollars per million tokens:

$ per 1M tokensGemini 3.8 Flash (to 31 Dec 2026)Gemini 3.8 Flash (from 1 Jan 2027)DeepSeek V4 Flash, peakDeepSeek V4 Flash, off-peak
Input0.751.500.440.22
Cached input0.0750.150.0140.007
Output3.757.501.320.66

DeepSeek’s peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday; everything else is off-peak at exactly half. That window is seven hours a day on weekdays, so a batch job you can schedule will almost never pay the peak column. These are the rates announced in DeepSeek’s changelog entry of 13 August 2026 and effective from 16:00 UTC on 16 August, when DeepSeek said it would “adopt peak/off-peak pricing, with off-peak prices set at half of the peak-hour prices”. Before that date V4 Flash listed at $0.14 input and $0.28 output, so even the off-peak column is a 1.6x to 2.4x increase on launch pricing.

Now the bill. Artificial Analysis paid $825.83 to run Gemini 3.8 Flash (high) through the full Intelligence Index and $323.26 for DeepSeek V4 Flash 0731, at $0.44 / $1.32, the peak rate. On the same task set, that is a 2.6x gap for seven points.

AA also prints a Cost per Task figure, computed separately from the index total: $0.58 high, $0.41 medium and $0.24 low for Gemini 3.8 Flash, and $0.11 for DeepSeek V4 Flash 0731 at its $0.44 / $1.32 peak rate. Off-peak halves every DeepSeek line, so roughly $0.055; that last number is our arithmetic, not one AA prints. The two AA measures do not agree on the ratio, 2.6x on the index total against 5.3x on Cost per Task, so quote whichever matches how you buy: whole-workload totals, or per-call averages.

Put the 52-point models side by side:

At AA score 52Cost per task
Gemini 3.8 Flash (low)$0.24
DeepSeek V4 Flash 0731, peak or peak-flat host$0.11 (AA)
DeepSeek V4 Flash 0731, off-peak~$0.055 (derived)

At the same intelligence, DeepSeek costs about half per task even at peak rates, and about a quarter off-peak. What Gemini 3.8 Flash at low effort buys for the extra $0.13 is output about 2.2x faster and image, audio, video and file input.

Caching and the 2027 Price Increase

Caching. Gemini 3.8 Flash reads cached input at $0.075, a tenth of fresh input. DeepSeek reads a cache hit at $0.014 peak, one thirty-first of a miss. If your prompts share a long prefix, DeepSeek’s cache is the cheapest input token on either price list. If they do not, the column is irrelevant.

The 2027 cliff. On 1 January 2027 Gemini 3.8 Flash’s rates double, so its $0.58 / $0.41 / $0.24 per-task ladder becomes roughly $1.16 / $0.82 / $0.48 if token counts hold. DeepSeek’s pricing page commits to nothing beyond “Product prices may vary and DeepSeek reserves the right to adjust them.” Any twelve-month cost model needs both of those rows.

When to Pick Gemini 3.8 Flash

  • The task needs more than a 53. There is no DeepSeek option above that line.
  • Inputs include images, audio, video or PDFs.
  • Latency matters: ~302 tokens per second, and AA’s Time per Task of 0.8 minutes at low effort against 2.5 minutes at high.
  • You cannot schedule your traffic around a seven-hour UTC window.

When to Pick DeepSeek V4 Flash 0731

  • A 52 is enough and you can run off-peak, or your prompts share long cached prefixes.
  • You need 384K output tokens in one call.
  • You want open weights you can host yourself later.
  • Text only is fine. The vision build is a separate, experimental model.

When Neither Is the Right Answer

If you are already on Gemini 3.7 Flash and it does the job, Google itself says to stay. 3.7 Flash scores 56 for $0.40 per AA task; 3.8 Flash at medium scores 57 for $0.41. That one point is close to free, but moving to high effort for 59 is not, and 3.7 Flash remains fully supported. See the Gemini 3.7 Flash API guide for the effort-level pricing detail.

If you were choosing between DeepSeek V4 Flash and Gemini 3.6 Flash, the earlier comparison still holds at the 52 level; 3.8 Flash at low effort joins that tie at its own $0.75 / $3.75 and, per AA, about 30% lower cost per task than 3.6 Flash at high.

Try Both via Ofox: One Endpoint, One Key

DeepSeek V4 Flash 0731 is in the Ofox catalog as deepseek/deepseek-v4-flash-0731, at $0.44 input, $1.32 output and $0.014 cache read per million tokens. Those are the same figures as DeepSeek’s peak-hour rate, billed flat: there is no off-peak window on Ofox. Which means AA’s $0.11 per task, measured at exactly those rates, is the row that applies; the off-peak row does not.

Gemini 3.8 Flash is not in the Ofox catalog as of 3 September 2026. google/gemini-3.7-flash is, at $0.75 input and $3.75 output on its Google provider route (the CloudVertex route on the same model page is $1.50 / $7.50, which is what the pricing field in /v1/models currently reports), the same per-token price 3.8 Flash carries. The loop below runs DeepSeek against 3.7 Flash today; when 3.8 Flash lands, it is a one-string change. The live catalog at GET /v1/models is the authority on what you can call.

from openai import OpenAI

client = OpenAI(base_url="https://api.ofox.ai/v1", api_key="YOUR_OFOX_API_KEY")

MODELS = [
    "deepseek/deepseek-v4-flash-0731",
    "google/gemini-3.7-flash",   # swap for google/gemini-3.8-flash once it is listed
]

prompt = "Refactor this function to remove the nested loop. Return only code."

for model in MODELS:
    r = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
    )
    u = r.usage
    print(f"{model:36} in={u.prompt_tokens:6} out={u.completion_tokens:6}")

Log completion_tokens, not just latency. Everything in this post that is not a sticker price comes down to that column on your own traffic.

Sources

Prices, index scores and token counts checked on 3 September 2026. Ofox catalog listing checked against the live /v1/models endpoint the same day; Ofox prices are from the model pages.

Frequently Asked Questions

Is Gemini 3.8 Flash better than DeepSeek V4 Flash?
On the Artificial Analysis Intelligence Index, yes: Gemini 3.8 Flash scores 59 at high reasoning effort, 57 at the default medium, and 52 at low. DeepSeek V4 Flash 0731 scores 52 at max effort. Nothing in DeepSeek's current lineup reaches 59; V4 Pro 0813 scores 53. At low effort, Gemini 3.8 Flash and DeepSeek V4 Flash are tied on the index.
How much does Gemini 3.8 Flash cost compared with DeepSeek V4 Flash?
List price: Gemini 3.8 Flash is $0.75 per million input tokens and $3.75 output through 31 December 2026, then $1.50 and $7.50 from 1 January 2027. DeepSeek V4 Flash is $0.44 input (cache miss), $0.014 cache hit and $1.32 output at peak hours, and half that off-peak. Artificial Analysis paid $825.83 to run Gemini 3.8 Flash (high) through its full Intelligence Index and $323.26 for DeepSeek V4 Flash 0731, a 2.6x gap.
Why is the real cost gap smaller than the sticker gap?
On AA's index total the gap is 2.6x, below the 2.8x output-price ratio, because DeepSeek V4 Flash 0731 writes far more tokens: Artificial Analysis recorded 210M output tokens for DeepSeek across the index against 120M for Gemini 3.8 Flash, 1.75x as many. AA's separately computed Cost per Task widens the gap to 5.3x ($0.58 vs $0.11). Output speed widens the gap in the other direction: about 302 tokens per second for Gemini 3.8 Flash against 138 for DeepSeek.
Does Gemini 3.8 Flash use more tokens than 3.7 Flash?
Yes. Google says 3.8 Flash 'works harder' and 'might use more tokens to maximize performance, especially at higher effort levels'. Artificial Analysis measured a 30% increase in average output tokens per task, to 48k, and cost per index task rising from $0.40 on 3.7 Flash to $0.58 on 3.8 Flash at high effort, despite identical per-token prices.
Can I call both through Ofox today?
DeepSeek V4 Flash 0731 is in the Ofox catalog as deepseek/deepseek-v4-flash-0731 at $0.44 input, $1.32 output and $0.014 cache read per million, the same figures as DeepSeek's peak-hour rate, with no off-peak window. Gemini 3.8 Flash is not in the catalog as of 3 September 2026; google/gemini-3.7-flash is, at the same $0.75 / $3.75 as 3.8 Flash.