Qwen3.5-35B-A3B
qwen3.5-35b-a3bThe Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... Built by Qwen.
Prices updated
Input price
$0.15
per 1M tokens · Standard
Output price
$1.00
per 1M tokens · Standard
Input limit
262K
tokens
Output limit
236K
tokens
Input formats
Output formats
Qwen3.5-35B-A3B price history
4 price records since 9/25/2026 · 2 changes
Qwen3.5-35B-A3B cost calculator
$8.00/month
Standard pricing
Overview
What is Qwen3.5-35B-A3B?
Qwen3.5-35B-A3B is a text & reasoning and vision model from Qwen. The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall... Its 262K context window and $0.15 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
Qwen3.5-35B-A3B benchmarks & speed
How Qwen3.5-35B-A3B scores on standardized evaluations, and where it lands among every model we track.
19.3
Intelligence Index
145tok/s
Output speed
Reasoning
GPQA Diamond
84.5%
Graduate-level scientific reasoning · top 43%
HLE
21.0%
Humanity's Last Exam · top 50%
Latency & design
- Time to first token
- 2.18s
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Qwen3.5-35B-A3B pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.15
Per 1M tokens
Cached input · Standard
$0.16
Per 1M tokens
Capabilities
Qwen3.5-35B-A3B Tools
Tools available when using Qwen3.5-35B-A3B through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Qwen3.5-35B-A3B API providers
7 providers serve Qwen3.5-35B-A3B. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| Darkbloomfp4 | $0.08 | $0.75 | $0.04 | 262K | 99.68% |
| DeepInfrafp8 | $0.14 | $1.00 | $0.05 | 262K | 99.96% |
| Parasailfp8 | $0.15 | $1.00 | $0.05 | 262K | 99.95% |
| Alibaba | $0.16 | $1.30 | — | 262K | 86.85% |
| AtlasCloudfp8 | $0.23 | $1.80 | $0.23 | 262K | 99.77% |
| SiliconFlowfp8 | $0.24 | $1.80 | $0.15 | 262K | 98.24% |
| Venice | $0.31 | $1.25 | $0.16 | 256K | 99.84% |
- Input
- $0.08
- Output
- $0.75
- Cached
- $0.04
- Context
- 262K
- Uptime
- 99.68%
- Input
- $0.14
- Output
- $1.00
- Cached
- $0.05
- Context
- 262K
- Uptime
- 99.96%
- Input
- $0.15
- Output
- $1.00
- Cached
- $0.05
- Context
- 262K
- Uptime
- 99.95%
- Input
- $0.16
- Output
- $1.30
- Cached
- —
- Context
- 262K
- Uptime
- 86.85%
- Input
- $0.23
- Output
- $1.80
- Cached
- $0.23
- Context
- 262K
- Uptime
- 99.77%
- Input
- $0.24
- Output
- $1.80
- Cached
- $0.15
- Context
- 262K
- Uptime
- 98.24%
- Input
- $0.31
- Output
- $1.25
- Cached
- $0.16
- Context
- 256K
- Uptime
- 99.84%
6 of 7
Strengths and limitations
Strengths
- Low input cost at $0.15 per million tokens suits high-volume workloads.
- 262K context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Generated tokens cost 7× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
