Back to blog

Python · Tutorial · Transcripts

Get Transcripts for Every Video on a YouTube Channel or Playlist

List a channel's or playlist's videos page by page, then fetch the transcripts 100 at a time. Python code that resumes where it stopped, and what a 1,000-video channel costs.

· 2 min read · YTAPI

Pulling every transcript from a channel is two jobs: get the list of video IDs, then fetch a transcript for each one. The first is a handful of requests. The second is hundreds or thousands, and it's where scripts usually fall over: one video at a time is slow, and a long loop from a server tends to get blocked halfway through.

If you already use yt-dlp, yt-dlp --flat-playlist --print id "https://www.youtube.com/@TED/videos" prints a channel's video IDs without downloading anything. The rest of this post does both steps over HTTP with YTAPI, so there's nothing to install beyond requests.

Step 1: list the videos

GET /v1/channels/{id}/videos takes a handle (@TED) or a channel ID (UC...) and returns 30 videos per page, newest first by default. Each page has has_more and a next_cursor to pass back as cursor:

import os
import requests

API = "https://api.ytapi.dev"
HEADERS = {"Authorization": f"Bearer {os.environ['YTAPI_KEY']}"}

def channel_video_ids(channel: str, sort_by: str = "newest"):
    cursor = None
    while True:
        params = {"sort_by": sort_by}
        if cursor:
            params["cursor"] = cursor
        res = requests.get(f"{API}/v1/channels/{channel}/videos", headers=HEADERS, params=params, timeout=30)
        res.raise_for_status()
        page = res.json()
        for video in page["videos"]:
            yield video["video_id"]
        if not page.get("has_more") or not page.get("next_cursor"):
            return
        cursor = page["next_cursor"]

sort_by can also be popular or oldest. For a playlist, GET /v1/playlists/{id} pages the same way: the first page also carries the playlist's title and video_count, and each video has an index.

Each page costs 1 credit.

Step 2: fetch the transcripts in batches

POST /v1/batch takes up to 100 tasks and runs up to 20 of them at a time. It answers 202 with a job ID right away; GET /v1/batch/{id} is free to poll. Each result carries the HTTP status it would have had on its own, and either data or an error:

import pathlib
import time

OUT = pathlib.Path("transcripts")

def fetch_all(video_ids, fmt="text"):
    OUT.mkdir(exist_ok=True)
    todo = [v for v in video_ids if not (OUT / f"{v}.txt").exists()]  # resume after a crash
    for i in range(0, len(todo), 100):
        tasks = [{"id": v, "type": "transcript", "video_id": v, "format": fmt} for v in todo[i:i + 100]]
        job = requests.post(f"{API}/v1/batch", headers=HEADERS,
                            json={"tasks": tasks, "concurrency": 10}, timeout=30)
        job.raise_for_status()
        job_id = job.json()["id"]

        while True:
            time.sleep(2)
            status = requests.get(f"{API}/v1/batch/{job_id}", headers=HEADERS, timeout=30).json()
            if status["status"] in ("completed", "failed"):
                break

        for r in status.get("results", []):
            if r["status"] == 200:
                (OUT / f"{r['video_id']}.txt").write_text(r["data"], encoding="utf-8")
            else:
                print(r["video_id"], r["status"], r["error"]["code"])

fetch_all(list(channel_video_ids("@TED")))

The script skips videos it has already saved, so if it stops halfway you run it again and it continues. Videos that had no captions are asked for again on the next run, which costs nothing. Videos without captions come back as 404 with captions_disabled and aren't billed; private and removed videos return video_unavailable.

format can be text (as here), markdown for LLM input, srt or vtt for subtitle files, or segments for JSON with timestamps. With a JSON format, data is an object rather than a string, so save it with json.dumps instead. Without languages, each transcript comes back in the video's own language; pass ["en", "*"] in each task if you want English whenever a video has it.

What it costs

For a channel with 1,000 videos:

RequestsCredits
List the videos34 pages34
Transcripts10 batches of 1001 per video that has captions
Pollinga few per batchfree

So at most about 1,034 credits, less for every video without captions. A $9 pack of 2,000 covers it, and credits don't expire.

FAQ

How long does a 1,000-video channel take?

Listing takes a few seconds. Transcripts run 10 or 20 at a time inside each batch, and most come back in under a second, so a batch of 100 usually finishes well within a minute. Raise concurrency to 20 if you're in a hurry.

Can I do this without batches?

Yes. Loop over the IDs and call POST /v1/transcripts for each, with a small thread pool. Batches just save you from managing concurrency and retries yourself. The Python guide has the single-video version with retries.

What about Shorts and live streams?

YouTube keeps Shorts and live streams on their own channel tabs, so check that the list from step 1 has the videos you expect. Step 2 works for any public video ID, whatever kind of video it is.