Transcribe 1 Pro is a speech-to-text model from Fish Audio tuned for interviews, meetings, and podcasts. It labels speakers with inline speaker markers, preserves emotion and vocal-event cues such as [laughter], detects language automatically, and can return timestamped word-level segments.
| $0.0001 | 0.83s |
P50, best provider
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.