PriceIndex
QW

Qwen3.5-35B-A3B

qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... Built by Qwen.

Prices updated

Input price

$0.15

per 1M tokens · Standard

Output price

$1.00

per 1M tokens · Standard

Input limit

262K

tokens

Output limit

236K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

Qwen3.5-35B-A3B price history

4 price records since 9/25/2026 · 2 changes

Qwen3.5-35B-A3B cost calculator

Input20M tokens
Output5M tokens

$8.00/month

Standard pricing

    Overview

    What is Qwen3.5-35B-A3B?

    Qwen3.5-35B-A3B is a text & reasoning and vision model from Qwen. The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... Its 262K context window and $0.15 input price make it a candidate for cost-sensitive, high-throughput applications.

    Image understandingDocument analysisContent workflowsClassification

    Benchmarks

    Qwen3.5-35B-A3B benchmarks & speed

    How Qwen3.5-35B-A3B scores on standardized evaluations, and where it lands among every model we track.

    19.3

    Intelligence Index

    Better than 49% of 153 models

    145tok/s

    Output speed

    Better than 75% of 119 models

    Reasoning

    GPQA Diamond

    84.5%

    Graduate-level scientific reasoning · top 43%

    HLE

    21.0%

    Humanity's Last Exam · top 50%

    Latency & design

    Time to first token
    2.18s

    Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.

    Rates

    Qwen3.5-35B-A3B pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $0.15

    Per 1M tokens

    Cached input · Standard

    $0.16

    Per 1M tokens

    Capabilities

    Qwen3.5-35B-A3B Tools

    Tools available when using Qwen3.5-35B-A3B through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Supported

    JSON mode

    Supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Supported

    Where to run it

    Qwen3.5-35B-A3B API providers

    7 providers serve Qwen3.5-35B-A3B. Prices are per 1M tokens; uptime is the last 24 hours.

    Darkbloomfp4
    Input
    $0.08
    Output
    $0.75
    Cached
    $0.04
    Context
    262K
    Uptime
    99.68%
    DeepInfrafp8
    Input
    $0.14
    Output
    $1.00
    Cached
    $0.05
    Context
    262K
    Uptime
    99.96%
    Parasailfp8
    Input
    $0.15
    Output
    $1.00
    Cached
    $0.05
    Context
    262K
    Uptime
    99.95%
    Alibaba
    Input
    $0.16
    Output
    $1.30
    Cached
    —
    Context
    262K
    Uptime
    86.85%
    AtlasCloudfp8
    Input
    $0.23
    Output
    $1.80
    Cached
    $0.23
    Context
    262K
    Uptime
    99.77%
    SiliconFlowfp8
    Input
    $0.24
    Output
    $1.80
    Cached
    $0.15
    Context
    262K
    Uptime
    98.24%

    6 of 7

    Strengths and limitations

    Strengths

    • Low input cost at $0.15 per million tokens suits high-volume workloads.
    • 262K context supports large documents, repositories and extended conversations.
    • Supports text & reasoning and vision workloads in one model.

    Limitations

    • Generated tokens cost 7× more than input tokens, which matters for verbose responses.
    • Large context capacity does not guarantee consistent retrieval across the entire prompt.
    • Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from Qwen

    View provider →

    $0.36 in · $0.40 out

    View model

    $0.10 in · $0.20 out

    View model

    $0.66 in · $1.00 out

    View model
    1M

    $0.26 in · $0.78 out

    View model

    Frequently asked questions about Qwen3.5-35B-A3B

    Qwen3.5-35B-A3B is a text & reasoning and vision model from Qwen. The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...