[Alibaba Cloud](/ai-gateway/models/labs/alibaba)

# Qwen3 Embedding 8B

Qwen3 Embedding 8B is Alibaba Cloud's 8B-tier text embedding model in the Qwen3 Embedding line, producing 4096-dimensional vectors and ranking first on the MTEB multilingual leaderboard at release, built for demanding cross-lingual retrieval and RAG workloads. Your use is subject to Alibaba Cloud's [Terms](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-product-terms-of-service-v-3-8-0) & [Privacy](https://www.alibabacloud.com/help/en/legal/latest/alibaba-cloud-international-website-privacy-policy) Policies.

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

```
1import { embed } from 'ai';
2

3const result = await embed({
4  model: 'alibaba/qwen3-embedding-8b',
5  value: 'Sunny day at the beach',
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/qwen3-embedding-8b) [About](/ai-gateway/models/qwen3-embedding-8b/about) [Providers](/ai-gateway/models/qwen3-embedding-8b/providers) [Similar](/ai-gateway/models/qwen3-embedding-8b/similar) [FAQ](/ai-gateway/models/qwen3-embedding-8b/faq)

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Input | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- |

| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 33K | $0.05/M |  |  |  | 06/05/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- |
| ![nebius logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnebius.png&w=48&q=75) [Nebius](/ai-gateway/models/providers/nebius) Legal:[Terms](https://docs.nebius.com/legal/terms-of-use?_gl=1*17d22j*_gcl_au*MTAwNjQ2OTE2LjE3NjYwOTcyMzcuMTI0OTg1NTgzNi4xNzY2MDk3ODgzLjE3NjYwOTg1MzY.)•[Privacy](https://docs.nebius.com/legal/privacy?_gl=1*17d22j*_gcl_au*MTAwNjQ2OTE2LjE3NjYwOTcyMzcuMTI0OTg1NTgzNi4xNzY2MDk3ODgzLjE3NjYwOTg1MzY.) | 41K | $0.01/M |  |  |  | 06/05/2025 |  |

## [Copy link to heading](#more-models-by-alibaba-cloud)More models by Alibaba Cloud

All

Text

Code

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-27b](/ai-gateway/models/qwen3.8-27b) | 1M | 0.3s | 55tps | $0.10/M | $0.40/M | Read:$0.01/M Write:— | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![runinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fruninfra.png%3Fv%3D1786910657446&w=48&q=75) |  |  | 08/14/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-2.4t-a95b](/ai-gateway/models/qwen3.8-2.4t-a95b) | 262K | 1.3s | 182tps | $1.65/M | $4.95/M | Read:$0.12/M Write:$2.50/M | — |  | ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![gmicloud logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fgmicloud.png%3Fv%3D1785888489595&w=48&q=75) +3 |  |  | 08/03/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.8-max](/ai-gateway/models/qwen3.8-max) | 1M | 4.2s | 48tps | $2/M | $6/M | Read:$0.25/M Write:$2.50/M | — | +1 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  | 08/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-flash](/ai-gateway/models/qwen3.7-flash) | 991K | 3.7s | 113tps | $0.03/M+2 more | $0.13/M+2 more | Read: $0.006/M+2 more Write: $0.04/M+2 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  | 07/28/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-plus](/ai-gateway/models/qwen3.7-plus) | 1M | 2.0s | 349tps | $0.40/M+1 more | $1.60/M+1 more | Read: $0.08/M+1 more Write: $0.50/M+1 more | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) |  |  | 06/02/2026 |  |
| ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%20model%20logo.png&w=48&q=75) [alibaba/qwen3.7-max](/ai-gateway/models/qwen3.7-max) | 991K | 2.8s | 56tps | $2.50/M | $7.50/M | Read:$0.50/M Write:$3.13/M | — |  | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) |  |  | 05/21/2026 |  |

## [Copy link to heading](#about-qwen3-embedding-8b)About Qwen3 Embedding 8B

Qwen3 Embedding 8B is purpose-built for retrieval workloads where accuracy is paramount. Qwen3 Embedding 8B ranked first on the MTEB multilingual leaderboard with a score of 70.58 at release. Its 4096-dimensional output space encodes fine-grained semantic distinctions that smaller embedding models flatten.

The architecture employs 36 transformer layers and derives from the Qwen3 foundation model. The resulting embeddings generalize across both in-domain and out-of-domain retrieval scenarios.

Coverage spans more than 100 natural languages and multiple programming languages, enabling truly multilingual vector indexes where documents in French, Japanese, or Python code can be searched using queries in any supported language. Matryoshka Representation Learning lets operators shorten vectors at inference time, helpful for tiered index architectures where a coarse first-pass retrieval uses short vectors and a reranking stage uses full-resolution representations.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: For workloads indexing sensitive documents, confirm that your chosen provider's data-residency region aligns with your compliance requirements before routing production traffic.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-qwen3-embedding-8b)When to Use Qwen3 Embedding 8B

### Best for

- MTEB-driven production retrieval: Systems where MTEB multilingual scores from the model's release evaluations are the primary criterion
- Long-document RAG: Pipelines that benefit from context of 41K tokens and 4096-dimensional representations to preserve semantic detail
- Cross-lingual knowledge bases: Indexes spanning many languages and programming environments
- Research and evaluation workloads: MTEB-adjacent benchmarks serve as a proxy for real retrieval performance

### Consider alternatives when

- Tight embedding cost budgets: Per-token cost dominates and slightly lower accuracy is acceptable, the 0.6B or 4B variants may provide sufficient quality
- Memory-constrained deployments: Environments with strict memory limits make a fully loaded 8B model impractical
- Generative output required: This model produces embeddings only; use a generative model when you need text output

## [Copy link to heading](#conclusion)Conclusion

Qwen3 Embedding 8B is the right tool when retrieval accuracy across languages and domains can't be compromised. Its first-place standing on the MTEB multilingual leaderboard at release and its 4096-dimensional output make it a strong foundation for enterprise-grade semantic search and RAG systems willing to invest in embedding quality.