DeepSeek V4.1 Flash

Chat
deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is DeepSeek's latest efficiency-optimized Mixture-of-Experts model with a 1M-token context window and up to 384K output tokens. It adds native multimodal vision understanding, supports thinking mode (on by default), tool calls, JSON output and prompt caching, and per DeepSeek surpasses V4 Pro on quality, cost and speed. Served upstream under the official model id deepseek-flash.

Fenêtre de contexte
1M
Tokens de sortie max
384K
Publié
2026-09-10
Capacités
VisionFunction CallingRaisonnementPrompt Caching
Fournisseurs disponibles
DeepSeek
Protocoles supportés
openaianthropic

Providers

DeepSeek
Tokens d'entrée
$0.3/M
Tokens de sortie
$1.2/M
Lecture cache
$0.006/M
Protocols
openai/v1/chat/completions/v1/responses
anthropic

Exemples de code

from openai import OpenAI
client = OpenAI(
base_url="https://api.ofox.ai/v1",
api_key="YOUR_OFOX_API_KEY",
)
response = client.chat.completions.create(
model="deepseek/deepseek-v4.1-flash",
messages=[
{"role": "user", "content": "Hello!"}
],
)
print(response.choices[0].message.content)

Questions fréquentes

DeepSeek V4.1 Flash sur Ofox.ai coûte $0.3/M par million de tokens d'entrée et $1.2/M par million de tokens de sortie. Paiement à l'usage, sans frais mensuels.