[Fish Audio](/ai-gateway/models/labs/fish-audio)

# Transcribe-1

Transcribe-1 is Fish Audio's multilingual speech-to-text model, with automatic language detection and word-level timestamps for recorded audio. Your use is subject to Fish Audio's [Terms](https://fish.audio/terms/) & [Privacy](https://fish.audio/privacy/) Policies.

Free

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

```
1import { experimental_transcribe as transcribe } from 'ai';
2import { gateway } from '@ai-sdk/gateway';
3import { readFile } from 'node:fs/promises';
4

5const result = await transcribe({
6  model: gateway.transcriptionModel('fish-audio/transcribe-1'),
7  audio: await readFile('audio.mp3'),
8});
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/transcribe-1) [About](/ai-gateway/models/transcribe-1/about) [Providers](/ai-gateway/models/transcribe-1/providers) [Similar](/ai-gateway/models/transcribe-1/similar) [FAQ](/ai-gateway/models/transcribe-1/faq)

## [Copy link to heading](#playground)Playground

Try out Transcribe-1 by Fish Audio. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75)Transcribe-1

### Speech to text

Record a short clip from your microphone and the model transcribes it to text.

Idle

![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=96&q=75)

Record a clip to see the transcript here.

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Input | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- |

| ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) [Fish Audio](/ai-gateway/models/providers/fish-audio) Free Legal:[Terms](https://fish.audio/terms/)•[Privacy](https://fish.audio/privacy/) | Free |  |  |  | 03/01/2026 |  |
| --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#more-models-by-fish-audio)More models by Fish Audio

All

Speech

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) [fish-audio/s2.1-pro](/ai-gateway/models/s2.1-pro) |  |  |  | Free | Free |  | — |  | ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) |  |  | 07/28/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) [fish-audio/s2-pro](/ai-gateway/models/s2-pro) |  |  |  | Free | Free |  | — |  | ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) |  |  | 03/09/2026 |  |
| ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) [fish-audio/s1](/ai-gateway/models/s1) |  |  |  | Free | Free |  | — |  | ![fish-audio logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffish-audio.png%3Fv%3D1785975667162&w=48&q=75) |  |  | 10/20/2025 |  |

## [Copy link to heading](#about-transcribe-1)About Transcribe-1

Transcribe-1 is Fish Audio's multilingual speech-to-text model, and the counterpart to the S-series voice models in the same family. Where those generate speech, Transcribe-1 transcribes it.

Language detection is automatic, so you do not have to declare the spoken language ahead of time. That matters for mixed-language archives and for user-generated audio where the language is not known in advance.

When alignment detail is requested, Transcribe-1 returns word-level timestamped segments rather than a single block of text. That is what subtitle generation, transcript search, and clip extraction all depend on, since each needs to map words back to positions in the audio.

Audio is supplied either as a URL or as base64. Speech-to-text through AI Gateway is billed by audio duration rather than by tokens, so cost tracks the length of the recording. See the pricing panel on this page for current rates.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Fish Audio publishes less public documentation for Transcribe-1 than for its voice models, and its own models overview covers the speech generation family rather than transcription. Verify supported languages, file size limits, and duration caps against your own test audio before committing a pipeline to it.
- Configuration: Billing follows audio duration, not tokens. A long recording costs the same whether it is dense speech or mostly silence, so trimming leading and trailing silence before upload is worth doing on large batches.
- Configuration: Transcribe-1 handles recorded audio. A live conversational surface needs a realtime speech model rather than a transcription pass over a finished file.
- Zero Data Retention: Zero Data Retention is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-transcribe-1)When to Use Transcribe-1

### Best for

- Recorded Audio Transcription: Calls, meetings, podcasts, and voice notes
- Mixed-Language Archives: Automatic detection without declaring a language
- Subtitle Generation: Word-level timestamped segments
- Transcript Search: Matched words mapped back to audio positions
- Clip Extraction: Driven by word timing rather than manual scrubbing

### Consider alternatives when

- Live Conversation: A realtime speech model fits a live surface better
- Documented Limits: Public documentation for this model is thin
- Speech Generation: The S-series models in this family produce audio
- Guaranteed Language Support: The supported set is not publicly enumerated

## [Copy link to heading](#conclusion)Conclusion

Transcribe-1 transcribes recorded audio with automatic language detection and word-level timing, which is what subtitles, transcript search, and clip extraction need. Point `fish-audio/transcribe-1` at AI Gateway for batch transcription, and validate language and duration limits against your own audio first.