Python · Tutorial · Transcripts
Get Transcripts for Every Video on a YouTube Channel or Playlist
List a channel's or playlist's videos page by page, then fetch the transcripts 100 at a time. Python code that resumes where it stopped, and what a 1,000-video channel costs.
· 2 min read · YTAPI
Pulling every transcript from a channel is two jobs: get the list of video IDs, then fetch a transcript for each one. The first is a handful of requests. The second is hundreds or thousands, and it's where scripts usually fall over: one video at a time is slow, and a long loop from a server tends to get blocked halfway through.
If you already use yt-dlp, yt-dlp --flat-playlist --print id "https://www.youtube.com/@TED/videos" prints a channel's video IDs without downloading anything. The rest of this post does both steps over HTTP with YTAPI, so there's nothing to install beyond requests.
Step 1: list the videos
GET /v1/channels/{id}/videos takes a handle (@TED) or a channel ID (UC...) and returns 30 videos per page, newest first by default. Each page has has_more and a next_cursor to pass back as cursor:
import os
import requests
API = "https://api.ytapi.dev"
HEADERS = {"Authorization": f"Bearer {os.environ['YTAPI_KEY']}"}
def channel_video_ids(channel: str, sort_by: str = "newest"):
cursor = None
while True:
params = {"sort_by": sort_by}
if cursor:
params["cursor"] = cursor
res = requests.get(f"{API}/v1/channels/{channel}/videos", headers=HEADERS, params=params, timeout=30)
res.raise_for_status()
page = res.json()
for video in page["videos"]:
yield video["video_id"]
if not page.get("has_more") or not page.get("next_cursor"):
return
cursor = page["next_cursor"]sort_by can also be popular or oldest. For a playlist, GET /v1/playlists/{id} pages the same way: the first page also carries the playlist's title and video_count, and each video has an index.
Each page costs 1 credit.
Step 2: fetch the transcripts in batches
POST /v1/batch takes up to 100 tasks and runs up to 20 of them at a time. It answers 202 with a job ID right away; GET /v1/batch/{id} is free to poll. Each result carries the HTTP status it would have had on its own, and either data or an error:
import pathlib
import time
OUT = pathlib.Path("transcripts")
def fetch_all(video_ids, fmt="text"):
OUT.mkdir(exist_ok=True)
todo = [v for v in video_ids if not (OUT / f"{v}.txt").exists()] # resume after a crash
for i in range(0, len(todo), 100):
tasks = [{"id": v, "type": "transcript", "video_id": v, "format": fmt} for v in todo[i:i + 100]]
job = requests.post(f"{API}/v1/batch", headers=HEADERS,
json={"tasks": tasks, "concurrency": 10}, timeout=30)
job.raise_for_status()
job_id = job.json()["id"]
while True:
time.sleep(2)
status = requests.get(f"{API}/v1/batch/{job_id}", headers=HEADERS, timeout=30).json()
if status["status"] in ("completed", "failed"):
break
for r in status.get("results", []):
if r["status"] == 200:
(OUT / f"{r['video_id']}.txt").write_text(r["data"], encoding="utf-8")
else:
print(r["video_id"], r["status"], r["error"]["code"])
fetch_all(list(channel_video_ids("@TED")))The script skips videos it has already saved, so if it stops halfway you run it again and it continues. Videos that had no captions are asked for again on the next run, which costs nothing. Videos without captions come back as 404 with captions_disabled and aren't billed; private and removed videos return video_unavailable.
format can be text (as here), markdown for LLM input, srt or vtt for subtitle files, or segments for JSON with timestamps. With a JSON format, data is an object rather than a string, so save it with json.dumps instead. Without languages, each transcript comes back in the video's own language; pass ["en", "*"] in each task if you want English whenever a video has it.
What it costs
For a channel with 1,000 videos:
| Requests | Credits | |
|---|---|---|
| List the videos | 34 pages | 34 |
| Transcripts | 10 batches of 100 | 1 per video that has captions |
| Polling | a few per batch | free |
So at most about 1,034 credits, less for every video without captions. A $9 pack of 2,000 covers it, and credits don't expire.
FAQ
How long does a 1,000-video channel take?
Listing takes a few seconds. Transcripts run 10 or 20 at a time inside each batch, and most come back in under a second, so a batch of 100 usually finishes well within a minute. Raise concurrency to 20 if you're in a hurry.
Can I do this without batches?
Yes. Loop over the IDs and call POST /v1/transcripts for each, with a small thread pool. Batches just save you from managing concurrency and retries yourself. The Python guide has the single-video version with retries.
What about Shorts and live streams?
YouTube keeps Shorts and live streams on their own channel tabs, so check that the list from step 1 has the videos you expect. Step 2 works for any public video ID, whatever kind of video it is.