Back to blog

Python · Tutorial · Transcripts

How to Get YouTube Transcripts in Python

Two ways to pull captions from YouTube in Python: the open-source youtube-transcript-api library, and a hosted HTTP API. Working code for both, and where each one breaks.

· 4 min read · YTAPI

There are two practical ways to get a YouTube transcript from Python. You can call YouTube's caption endpoints yourself with the open-source youtube-transcript-api library, or you can send one HTTP request to a hosted API that does that work for you. Both are a few lines of code. The difference shows up later, when the script leaves your laptop.

This post walks through both, with code you can paste, and ends with a short guide to choosing.

Working in JavaScript? The same two options are in the Node.js guide.

Option 1: the youtube-transcript-api library

youtube-transcript-api is the most widely used Python package for this. It needs no API key and no Google Cloud project.

pip install youtube-transcript-api
from youtube_transcript_api import YouTubeTranscriptApi

api = YouTubeTranscriptApi()
transcript = api.fetch("dQw4w9WgXcQ", languages=["en"])

for snippet in transcript:
    print(f"{snippet.start:7.2f}s  {snippet.text}")

Each snippet has text, start and duration in seconds. languages is a priority list, so ["de", "en"] tries German first and falls back to English. The library's API changed at version 1.0, so if you find older examples that call YouTubeTranscriptApi.get_transcript(...) as a static method, they were written for the old version.

Where it breaks

On a laptop or a home connection, the library usually just works. Two things tend to go wrong once it runs somewhere else:

  1. Blocked cloud IPs. Requests from AWS, Google Cloud, Azure and most other hosting providers are often refused by YouTube. The library raises RequestBlocked or IpBlocked. The fix it documents is to route traffic through rotating residential proxies, which is a separate service to buy and keep healthy. We wrote more about this in youtube-transcript-api blocked on AWS.
  2. Videos without the track you asked for. TranscriptsDisabled means the creator turned captions off; NoTranscriptFound means none of your languages exist. CouldNotRetrieveTranscript is the base class, and an IP block is a subclass of it, so catching only the base class hides which one you hit. New uploads may not have auto-generated captions yet. The three names are split out in TranscriptsDisabled, NoTranscriptFound, and CouldNotRetrieveTranscript, and the no-track case is in what to do when a video has no captions.

If you run a handful of videos from your own machine, neither will bother you much. If you run a job on a server every hour, the first one will.

Option 2: a hosted transcript API

A hosted API moves the YouTube side (proxies, retries, parsing) off your machine. You send a video ID and get JSON back. The example uses YTAPI, but the shape is similar for other providers.

import os
import requests

API_KEY = os.environ["YTAPI_KEY"]

def get_transcript(video_id: str) -> dict:
    res = requests.post(
        "https://api.ytapi.dev/v1/transcripts",
        headers={"Authorization": f"Bearer {API_KEY}"},
        json={"video_id": video_id, "format": "segments"},
        timeout=30,
    )
    if res.status_code == 404:
        # captions_disabled, language_not_found or video_unavailable
        raise LookupError(res.json()["error"]["code"])
    res.raise_for_status()
    return res.json()

data = get_transcript("dQw4w9WgXcQ")
print(data["language"], data["track_kind"])
for seg in data["segments"]:
    print(f"{seg['start']:7.2f}s  {seg['text']}")

video_id also accepts a full YouTube URL. track_kind tells you whether you got the creator's captions (manual) or YouTube's speech recognition (asr).

Languages

Without languages, the API returns the video's own language, the one spoken in it. The language field says which one you got. To choose, pass a list; "*" at the end means "fall back to the original language":

json={"video_id": video_id, "languages": ["es", "*"], "format": "segments"}

Other formats

format can also be markdown, text, srt or vtt. Those come back as the raw document rather than JSON, so read res.text:

res = requests.post(
    "https://api.ytapi.dev/v1/transcripts",
    headers={"Authorization": f"Bearer {API_KEY}"},
    json={"video_id": "dQw4w9WgXcQ", "format": "srt"},
    timeout=30,
)
open("captions.srt", "w").write(res.text)

For word-level timing on auto-generated tracks, use "format": "word_timestamps"; each segment then carries a words list with start and end times.

Fetching many videos

A thread pool is enough for most jobs. Keep concurrency modest and back off on 429, which is what any API returns when you exceed your rate limit:

import time
from concurrent.futures import ThreadPoolExecutor

def fetch_with_retry(video_id: str, attempts: int = 3):
    for i in range(attempts):
        res = requests.post(
            "https://api.ytapi.dev/v1/transcripts",
            headers={"Authorization": f"Bearer {API_KEY}"},
            json={"video_id": video_id, "format": "text"},
            timeout=30,
        )
        if res.status_code == 429:
            time.sleep(float(res.headers.get("Retry-After", 2 ** i)))
            continue
        if res.status_code == 404:
            return video_id, None
        res.raise_for_status()
        return video_id, res.text
    return video_id, None

ids = ["dQw4w9WgXcQ", "kCc8FmEb1nY", "jNQXAC9IVRw"]
with ThreadPoolExecutor(max_workers=5) as pool:
    for video_id, text in pool.map(fetch_with_retry, ids):
        print(video_id, "no captions" if text is None else f"{len(text)} chars")

For more than a few hundred videos at once, the batch endpoint takes up to 100 video IDs in one job; see the batch docs.

Which one to use

youtube-transcript-apiHosted API
CostFreePer successful request
Runs well onYour own machineAnywhere, including cloud servers
Blocked cloud IPsYour problem (proxies)The provider's problem
Setuppip installpip install requests and an API key
When YouTube changes somethingWait for a library releaseThe provider updates on their side

Start with the library if you are exploring on your own machine or processing a small, one-off list. Move to a hosted API when the code runs on a server, on a schedule, or for users who expect it to work every time.

If you want numbers before deciding: in our benchmark of recent, uncached videos from four regions, YTAPI answered every request on the first try with a median of 644 ms. You get 200 free credits on signup to try it on your own videos, and only HTTP 200 responses use a credit. After that, credit packs start at $9 for 2,000 and don't expire.

FAQ

Is youtube-transcript-api free?

Yes, it is open source and free to use. The cost shows up later: once your code runs on a cloud server, YouTube starts blocking the requests, and the usual fixes (residential proxies, or keeping a machine at home running) cost money or time. Why cloud servers get blocked.

How do I get transcripts for every video on a channel?

List the channel's video IDs first, with the channel videos endpoint (or the playlist endpoint for a playlist), then fetch the transcripts with the loop above or as batch jobs of up to 100 videos. The full script, with paging and resuming, is in transcripts for a whole channel or playlist.

Why do I get NoTranscriptFound when the video has captions?

Usually a language mismatch: the video has captions, but not in the language you asked for, or only auto-generated ones. Pass several languages, or "*" as a fallback. The error-by-error breakdown is in TranscriptsDisabled, NoTranscriptFound, and CouldNotRetrieveTranscript.