- Input /1M
- $1.25
- Output /1M
- $4.25
- Context
- 1M
- AA Intelligence
- 48.1
- DesignArena Elo
- 1358
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
Inference host
AI21 API pricing & models
AI21 Labs is an Israeli company specializing in natural language processing (NLP), developing advanced AI systems and foundation models designed for enterprise applications. Their mission is to empower businesses with state-of-the-art Large Language Models (LLMs) and AI applications to redefine how humans interact with text. AI21 serves open-weight models from other labs through its own API.
Overview
About AI21
AI21 Labs is an Israeli company specializing in natural language processing (NLP), developing advanced AI systems and foundation models designed for enterprise applications. Their mission is to empower businesses with state-of-the-art Large Language Models (LLMs) and AI applications to redefine how humans interact with text.
The company offers a suite of models and tools, including the Jurassic-2 family and the Jamba family. The Jamba models are particularly notable for their hybrid Mamba State Space Model (SSM) and Transformer architecture, which enables efficient processing of exceptionally long context windows.
AI21 Labs provides access to its models through its AI21 Studio developer platform, which includes a Python SDK and REST API, and also via major cloud providers like Amazon Bedrock and Google Cloud Vertex AI. They focus on solutions for high-value, data-intensive workflows such as grounded question answering, RAG workflows, and agentic systems, aiming to reduce hallucinations and improve accuracy for business-critical tasks.
Headquarters: ILPrivacy policyTermsStatus page
Strengths
- Jamba models offer an industry-leading context window of up to 256,000 tokens, enabling processing of massive documents for enterprises.
- The hybrid Mamba SSM and Transformer architecture of Jamba models provides high inference speed and efficiency, especially for long contexts.
- AI21 Labs provides specialized Task-Specific Models (TSMs) for common business needs like summarization, paraphrasing, and grammatical error correction.
- Maestro, an advanced AI system, helps deploy knowledge agents that automate data-intensive tasks, featuring RAG capabilities and self-validation to reduce hallu
- Models are available across multiple cloud platforms (AWS Bedrock, Google Cloud Vertex AI) and open-source hubs (Hugging Face, Kaggle) for flexible deployment.
- The platform supports advanced developer features like function calling and structured JSON output for integrating AI into complex applications.
Limitations
- Older model families, such as Jurassic-2 Ultra, have a significantly smaller context window (8,191 tokens) compared to the newer Jamba models.
- While Jamba's hybrid architecture is efficient, pure Transformer implementations can sometimes be faster for very short sequence lengths (below 2,000-4,000 toke
- Access to some models or higher rate limits may require contacting sales, indicating potential non-standardized access for specific enterprise needs.
Open-weight models
Models you can run on AI21
AI21 hosts popular open-weight models. Dedicated AI21 pricing is coming soon — here are the models it typically serves.
| $1.25 | $4.25 | 1M | 48.1 | 1358 | — | — | 0% | |
| $1.25 | $4.25 | 1M | 39.6 | 1318 | — | — | 0% | |
| $0.05 | $1.2 | 1M | 39.5 | 1325 | — | — | 357% | |
| $0.29232 | $0.58464 | 1M | 36 | 1244 | — | — | -19% | |
| $0.2156 | $0.6468 | 1M | 34.8 | — | — | — | 96% | |
| $0.03 | $0.056 | 1M | 34.3 | 1209 | — | — | 3% | |
| $1.25 | $4.25 | 1M | 33.7 | — | — | — | 0% | |
| $0.259 | $0.42 | 164K | 16 | — | — | — | 48% | |
| $0.27 | $0.41 | 164K | 16 | — | — | — | 0% | |
| $0.27 | $1 | 164K | 13.9 | — | — | — | 0% | |
| $0.25 | $0.95 | 164K | 13.7 | — | — | — | 0% | |
| $0.7 | $2.5 | 64K | 13.1 | — | 1239 | 80 | 3% | |
| $0.1875 | $0.6525 | 1M | 10 | 883 | — | — | 0% | |
| $0.29 | $1.14 | 164K | 9.7 | — | — | — | 167% | |
| $0.2574 | $1.0287 | 164K | 8.5 | — | 1211 | 75.4 | -20% | |
| $0.1 | $0.3 | 1.3M | 8.1 | 793 | — | — | 0% | |
| $0.8 | $0.8 | 8K | 7.9 | — | — | — | 0% | |
| $0.1 | $0.32 | 131K | 7.7 | — | — | — | 107% | |
| $0.05 | $0.08 | 131K | 6.9 | — | — | — | 0% | |
| $0.4 | $0.4 | 131K | 6.6 | — | — | — | 0% | |
| $0.05 | $0.33 | 131K | 5.7 | — | — | — | 0% | |
| $0.027 | $0.201 | 60K | 4.8 | — | — | — | 0% | |
| $0.018 | $0.32 | 1M | — | 1234 | — | — | 323% | |
| $0.66 | $1.98 | 1M | — | — | — | — | 282% | |
| $0.18 | $0.18 | 164K | — | — | — | — | 0% | |
| $0.3 | $1.2 | 131K | — | — | — | — | 133% | |
| $0.1 | $0.2 | 1M | — | — | — | — | 0% | |
| $0.1 | $0.2 | 1M | — | — | — | — | 0% | |
| $0.5 | $2.15 | 164K | — | — | — | — | 0% |
- Input /1M
- $1.25
- Output /1M
- $4.25
- Context
- 1M
- AA Intelligence
- 39.6
- DesignArena Elo
- 1318
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.05
- Output /1M
- $1.2
- Context
- 1M
- AA Intelligence
- 39.5
- DesignArena Elo
- 1325
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 357%
- Input /1M
- $0.29232
- Output /1M
- $0.58464
- Context
- 1M
- AA Intelligence
- 36
- DesignArena Elo
- 1244
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- -19%
- Input /1M
- $0.2156
- Output /1M
- $0.6468
- Context
- 1M
- AA Intelligence
- 34.8
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 96%
- Input /1M
- $0.03
- Output /1M
- $0.056
- Context
- 1M
- AA Intelligence
- 34.3
- DesignArena Elo
- 1209
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 3%
- Input /1M
- $1.25
- Output /1M
- $4.25
- Context
- 1M
- AA Intelligence
- 33.7
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.259
- Output /1M
- $0.42
- Context
- 164K
- AA Intelligence
- 16
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 48%
- Input /1M
- $0.27
- Output /1M
- $0.41
- Context
- 164K
- AA Intelligence
- 16
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.27
- Output /1M
- $1
- Context
- 164K
- AA Intelligence
- 13.9
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.25
- Output /1M
- $0.95
- Context
- 164K
- AA Intelligence
- 13.7
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.7
- Output /1M
- $2.5
- Context
- 64K
- AA Intelligence
- 13.1
- DesignArena Elo
- —
- Arena Elo
- 1239
- MMLU-Pro
- 80
- 30d trend
- 3%
- Input /1M
- $0.1875
- Output /1M
- $0.6525
- Context
- 1M
- AA Intelligence
- 10
- DesignArena Elo
- 883
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.29
- Output /1M
- $1.14
- Context
- 164K
- AA Intelligence
- 9.7
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 167%
- Input /1M
- $0.2574
- Output /1M
- $1.0287
- Context
- 164K
- AA Intelligence
- 8.5
- DesignArena Elo
- —
- Arena Elo
- 1211
- MMLU-Pro
- 75.4
- 30d trend
- -20%
- Input /1M
- $0.1
- Output /1M
- $0.3
- Context
- 1.3M
- AA Intelligence
- 8.1
- DesignArena Elo
- 793
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.8
- Output /1M
- $0.8
- Context
- 8K
- AA Intelligence
- 7.9
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.1
- Output /1M
- $0.32
- Context
- 131K
- AA Intelligence
- 7.7
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 107%
- Input /1M
- $0.05
- Output /1M
- $0.08
- Context
- 131K
- AA Intelligence
- 6.9
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.4
- Output /1M
- $0.4
- Context
- 131K
- AA Intelligence
- 6.6
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.05
- Output /1M
- $0.33
- Context
- 131K
- AA Intelligence
- 5.7
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.027
- Output /1M
- $0.201
- Context
- 60K
- AA Intelligence
- 4.8
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.018
- Output /1M
- $0.32
- Context
- 1M
- AA Intelligence
- —
- DesignArena Elo
- 1234
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 323%
- Input /1M
- $0.66
- Output /1M
- $1.98
- Context
- 1M
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 282%
- Input /1M
- $0.18
- Output /1M
- $0.18
- Context
- 164K
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.3
- Output /1M
- $1.2
- Context
- 131K
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 133%
- Input /1M
- $0.1
- Output /1M
- $0.2
- Context
- 1M
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.1
- Output /1M
- $0.2
- Context
- 1M
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
- Input /1M
- $0.5
- Output /1M
- $2.15
- Context
- 164K
- AA Intelligence
- —
- DesignArena Elo
- —
- Arena Elo
- —
- MMLU-Pro
- —
- 30d trend
- 0%
Showing 15 of 29
AI21 Labs' Jamba Models: Hybrid Architecture and Long Context
The Jamba family of models represents a significant offering from AI21 Labs, distinguishing itself with a unique hybrid Mamba State Space Model (SSM) and Transformer architecture. This innovative design combines the strengths of both paradigms: the linear scaling and efficiency of Mamba models for processing long sequences, and the strong performance and in-context learning capabilities of Transformers. As a result, Jamba models are capable of handling an extensive context window of up to 256,000 tokens, which is among the longest available for open models, making them highly suitable for applications requiring deep analysis of large documents, such as legal contracts, research papers, or financial reports.
Available in different sizes like Jamba 1.5 Mini and Jamba 1.5 Large, these models also incorporate a Mixture of Experts (MoE) component, which contributes to their inference efficiency. They are optimized for tasks like long-context Retrieval-Augmented Generation (RAG), grounded question answering, and data classification. Developers can access Jamba models through AI21 Studio, as well as managed services on Amazon Bedrock and Google Cloud Vertex AI, and can also deploy open-w
Enterprise Solutions and Task-Specific Capabilities
AI21 Labs focuses on providing robust AI solutions tailored for enterprise workflows. Beyond their foundational Jamba models, they offer AI21 Maestro, an advanced system designed to build and deploy intelligent agents that automate complex, data-intensive business tasks. Maestro features built-in RAG capabilities, semantic search, and web search integration, enabling agents to retrieve, reason, validate, and adapt information in real-time while adhering to performance and cost constraints. This system is specifically engineered to improve accuracy and significantly reduce AI hallucinations in business-critical applications.
For more specialized needs, AI21 Labs also provides a suite of Task-Specific Models (TSMs). These include Contextual Answers for accurate information retrieval from internal knowledge bases, Summarize for condensing lengthy texts, Paraphrase for rewriting content, and Grammatical Error Correction. These TSMs are designed to offer immediate value and higher accuracy for common commercial use cases, requiring less prompt engineering and allowing organizations to quickly integrate generative AI features into their products and services.
FAQ
AI21 pricing FAQ
Explore
