PriceIndex
TH

Inkling

inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... Built by Thinking Machines.

Prices updated

Input price

$1.00

per 1M tokens · Standard

Output price

$4.05

per 1M tokens · Standard

Input limit

524K

tokens

Output limit

472K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

Inkling price history

2 price records since 10/8/2026

Inkling cost calculator

Input20M tokens
Output5M tokens

$40/month

Standard pricing

    Overview

    What is Inkling?

    Inkling is a text & reasoning and vision model from Thinking Machines. Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,... Its 524K context window and $1.00 input price make it a candidate for quality-focused production applications.

    Image understandingDocument analysisContent workflowsClassification

    Rates

    Inkling pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $1.00

    Per 1M tokens

    Cached input · Standard

    $0.17

    Per 1M tokens

    Capabilities

    Inkling Tools

    Tools available when using Inkling through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Not supported

    JSON mode

    Not supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Not supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Supported

    Strengths and limitations

    Strengths

    • Priced at $1.00 input and $4.05 output per million tokens.
    • 524K context supports large documents, repositories and extended conversations.
    • Supports text & reasoning and vision workloads in one model.

    Limitations

    • Generated tokens cost 4× more than input tokens, which matters for verbose responses.
    • Large context capacity does not guarantee consistent retrieval across the entire prompt.
    • Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from Thinking Machines

    View provider →
    Inkling Small

    Thinking Machines

    524K

    $0.45 in · $1.20 out

    View model

    Frequently asked questions about Inkling

    Inkling is a text & reasoning and vision model from Thinking Machines. Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...