Python · Tutorial · Transcripts
How to Get YouTube Transcripts in Python
Two ways to pull captions from YouTube in Python: the open-source youtube-transcript-api library, and a hosted HTTP API. Working code for both, and where each one breaks.
· 4 min read · YTAPI
There are two practical ways to get a YouTube transcript from Python. You can call YouTube's caption endpoints yourself with the open-source youtube-transcript-api library, or you can send one HTTP request to a hosted API that does that work for you. Both are a few lines of code. The difference shows up later, when the script leaves your laptop.
This post walks through both, with code you can paste, and ends with a short guide to choosing.
Working in JavaScript? The same two options are in the Node.js guide.
Option 1: the youtube-transcript-api library
youtube-transcript-api is the most widely used Python package for this. It needs no API key and no Google Cloud project.
pip install youtube-transcript-apifrom youtube_transcript_api import YouTubeTranscriptApi
api = YouTubeTranscriptApi()
transcript = api.fetch("dQw4w9WgXcQ", languages=["en"])
for snippet in transcript:
print(f"{snippet.start:7.2f}s {snippet.text}")Each snippet has text, start and duration in seconds. languages is a priority list, so ["de", "en"] tries German first and falls back to English. The library's API changed at version 1.0, so if you find older examples that call YouTubeTranscriptApi.get_transcript(...) as a static method, they were written for the old version.
Where it breaks
On a laptop or a home connection, the library usually just works. Two things tend to go wrong once it runs somewhere else:
- Blocked cloud IPs. Requests from AWS, Google Cloud, Azure and most other hosting providers are often refused by YouTube. The library raises
RequestBlockedorIpBlocked. The fix it documents is to route traffic through rotating residential proxies, which is a separate service to buy and keep healthy. We wrote more about this in youtube-transcript-api blocked on AWS. - Videos without the track you asked for.
TranscriptsDisabledmeans the creator turned captions off;NoTranscriptFoundmeans none of your languages exist.CouldNotRetrieveTranscriptis the base class, and an IP block is a subclass of it, so catching only the base class hides which one you hit. New uploads may not have auto-generated captions yet. The three names are split out in TranscriptsDisabled, NoTranscriptFound, and CouldNotRetrieveTranscript, and the no-track case is in what to do when a video has no captions.
If you run a handful of videos from your own machine, neither will bother you much. If you run a job on a server every hour, the first one will.
Option 2: a hosted transcript API
A hosted API moves the YouTube side (proxies, retries, parsing) off your machine. You send a video ID and get JSON back. The example uses YTAPI, but the shape is similar for other providers.
import os
import requests
API_KEY = os.environ["YTAPI_KEY"]
def get_transcript(video_id: str) -> dict:
res = requests.post(
"https://api.ytapi.dev/v1/transcripts",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"video_id": video_id, "format": "segments"},
timeout=30,
)
if res.status_code == 404:
# captions_disabled, language_not_found or video_unavailable
raise LookupError(res.json()["error"]["code"])
res.raise_for_status()
return res.json()
data = get_transcript("dQw4w9WgXcQ")
print(data["language"], data["track_kind"])
for seg in data["segments"]:
print(f"{seg['start']:7.2f}s {seg['text']}")video_id also accepts a full YouTube URL. track_kind tells you whether you got the creator's captions (manual) or YouTube's speech recognition (asr).
Languages
Without languages, the API returns the video's own language, the one spoken in it. The language field says which one you got. To choose, pass a list; "*" at the end means "fall back to the original language":
json={"video_id": video_id, "languages": ["es", "*"], "format": "segments"}Other formats
format can also be markdown, text, srt or vtt. Those come back as the raw document rather than JSON, so read res.text:
res = requests.post(
"https://api.ytapi.dev/v1/transcripts",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"video_id": "dQw4w9WgXcQ", "format": "srt"},
timeout=30,
)
open("captions.srt", "w").write(res.text)For word-level timing on auto-generated tracks, use "format": "word_timestamps"; each segment then carries a words list with start and end times.
Fetching many videos
A thread pool is enough for most jobs. Keep concurrency modest and back off on 429, which is what any API returns when you exceed your rate limit:
import time
from concurrent.futures import ThreadPoolExecutor
def fetch_with_retry(video_id: str, attempts: int = 3):
for i in range(attempts):
res = requests.post(
"https://api.ytapi.dev/v1/transcripts",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"video_id": video_id, "format": "text"},
timeout=30,
)
if res.status_code == 429:
time.sleep(float(res.headers.get("Retry-After", 2 ** i)))
continue
if res.status_code == 404:
return video_id, None
res.raise_for_status()
return video_id, res.text
return video_id, None
ids = ["dQw4w9WgXcQ", "kCc8FmEb1nY", "jNQXAC9IVRw"]
with ThreadPoolExecutor(max_workers=5) as pool:
for video_id, text in pool.map(fetch_with_retry, ids):
print(video_id, "no captions" if text is None else f"{len(text)} chars")For more than a few hundred videos at once, the batch endpoint takes up to 100 video IDs in one job; see the batch docs.
Which one to use
| youtube-transcript-api | Hosted API | |
|---|---|---|
| Cost | Free | Per successful request |
| Runs well on | Your own machine | Anywhere, including cloud servers |
| Blocked cloud IPs | Your problem (proxies) | The provider's problem |
| Setup | pip install | pip install requests and an API key |
| When YouTube changes something | Wait for a library release | The provider updates on their side |
Start with the library if you are exploring on your own machine or processing a small, one-off list. Move to a hosted API when the code runs on a server, on a schedule, or for users who expect it to work every time.
If you want numbers before deciding: in our benchmark of recent, uncached videos from four regions, YTAPI answered every request on the first try with a median of 644 ms. You get 200 free credits on signup to try it on your own videos, and only HTTP 200 responses use a credit. After that, credit packs start at $9 for 2,000 and don't expire.
FAQ
Is youtube-transcript-api free?
Yes, it is open source and free to use. The cost shows up later: once your code runs on a cloud server, YouTube starts blocking the requests, and the usual fixes (residential proxies, or keeping a machine at home running) cost money or time. Why cloud servers get blocked.
How do I get transcripts for every video on a channel?
List the channel's video IDs first, with the channel videos endpoint (or the playlist endpoint for a playlist), then fetch the transcripts with the loop above or as batch jobs of up to 100 videos. The full script, with paging and resuming, is in transcripts for a whole channel or playlist.
Why do I get NoTranscriptFound when the video has captions?
Usually a language mismatch: the video has captions, but not in the language you asked for, or only auto-generated ones. Pass several languages, or "*" as a fallback. The error-by-error breakdown is in TranscriptsDisabled, NoTranscriptFound, and CouldNotRetrieveTranscript.