Python · Tutorial · Transcripts
Turn a YouTube Channel Into a Searchable Archive of Sermons, Lectures or Talks
Years of sermons, lectures or talks on one channel, and no way to find where something was said. A Python script that turns the channel into a full-text search with links to the exact minute, and keeps it up to date.
· 4 min read · YTAPI
A church that has streamed every Sunday for six years has about 300 sermons on YouTube. A university channel has a few hundred lectures; a podcast, a few hundred episodes. All of it was said on camera, and none of it can be searched. YouTube's search box looks at titles and descriptions, not at what was said in the videos.
This guide builds a small archive that fixes that. It downloads the transcript of every video on a channel, splits each one into short passages, and stores them in a SQLite full-text index. A search returns the passages that match, each with a link that opens the video at that moment. It runs on a laptop, and it costs about one credit per video once, then only the new uploads.
Channels like these are usually small, and their back catalogues aren't in anyone's cache. YTAPI fetches each video from YouTube when you ask, so it doesn't matter that nobody has requested them before. More on that in why transcript APIs slow down on small channels.
What you need
- Python 3.9 or later, with
requests(pip install requests). SQLite's full-text search, FTS5, ships with Python'ssqlite3on most systems. - A YTAPI key in the
YTAPI_KEYenvironment variable. New accounts get 200 free credits; a 300-video channel needs about 310. - The channel's handle, such as
@yourchurch, or itsUC...ID.
The script
Save this as archive.py, set CHANNEL, and run it. The first run takes a few minutes; later runs only fetch videos that are new.
import os
import sqlite3
import sys
import time
import requests
API = "https://api.ytapi.dev"
HEADERS = {"Authorization": f"Bearer {os.environ['YTAPI_KEY']}"}
CHANNEL = "@yourchannel"
db = sqlite3.connect("archive.db")
db.executescript("""
CREATE TABLE IF NOT EXISTS videos (
video_id TEXT PRIMARY KEY,
title TEXT,
status TEXT NOT NULL DEFAULT 'new' -- new, done or no_captions
);
CREATE VIRTUAL TABLE IF NOT EXISTS passages USING fts5(text, video_id UNINDEXED, start UNINDEXED);
""")
def list_new_videos(channel):
"""Newest first; stop at the first video already in the archive."""
found, cursor = [], None
while True:
params = {"sort_by": "newest"}
if cursor:
params["cursor"] = cursor
res = requests.get(f"{API}/v1/channels/{channel}/videos",
headers=HEADERS, params=params, timeout=30)
res.raise_for_status()
page = res.json()
for video in page["videos"]:
if db.execute("SELECT 1 FROM videos WHERE video_id = ?",
(video["video_id"],)).fetchone():
return found
found.append((video["video_id"], video.get("title", "")))
if not page.get("has_more") or not page.get("next_cursor"):
return found
cursor = page["next_cursor"]
def passages(segments, seconds=45):
"""Join caption segments into passages of about `seconds` each."""
chunk, start = [], None
for seg in segments:
if start is None:
start = seg["start"]
chunk.append(seg["text"])
if seg["end"] - start >= seconds:
yield start, " ".join(chunk)
chunk, start = [], None
if chunk:
yield start, " ".join(chunk)
def fetch_transcripts():
todo = [row[0] for row in db.execute("SELECT video_id FROM videos WHERE status = 'new'")]
for i in range(0, len(todo), 100):
tasks = [{"id": v, "type": "transcript", "video_id": v, "format": "segments"}
for v in todo[i:i + 100]]
job = requests.post(f"{API}/v1/batch", headers=HEADERS,
json={"tasks": tasks, "concurrency": 10}, timeout=30)
job.raise_for_status()
job_id = job.json()["id"]
while True:
time.sleep(2)
res = requests.get(f"{API}/v1/batch/{job_id}", headers=HEADERS, timeout=30)
res.raise_for_status()
status = res.json()
if status["status"] in ("completed", "failed"):
break
for r in status.get("results", []):
if r["status"] == 200:
for start, text in passages(r["data"]["segments"]):
db.execute("INSERT INTO passages (text, video_id, start) VALUES (?, ?, ?)",
(text, r["video_id"], int(start)))
db.execute("UPDATE videos SET status = 'done' WHERE video_id = ?", (r["video_id"],))
elif r["status"] == 404:
db.execute("UPDATE videos SET status = 'no_captions' WHERE video_id = ?",
(r["video_id"],))
db.commit()
print(f"{min(i + 100, len(todo))} of {len(todo)} videos fetched")
def search(query, limit=10):
rows = db.execute("""
SELECT videos.title, passages.video_id, passages.start,
snippet(passages, 0, '[', ']', '...', 12)
FROM passages JOIN videos ON videos.video_id = passages.video_id
WHERE passages MATCH ?
ORDER BY rank
LIMIT ?
""", (query, limit))
for title, video_id, start, snip in rows:
print(f"{title}\n https://youtu.be/{video_id}?t={start}\n {snip}\n")
if __name__ == "__main__":
if len(sys.argv) > 1:
search(" ".join(sys.argv[1:]))
else:
new = list_new_videos(CHANNEL)
db.executemany("INSERT OR IGNORE INTO videos (video_id, title) VALUES (?, ?)", new)
db.commit()
print(f"{len(new)} new videos")
fetch_transcripts()Then search:
python archive.py forgiveness
python archive.py '"fruit of the spirit"'
python archive.py 'grace NOT law'Each result is the video title, a link that starts the video at that passage, and the matching words in brackets. Searches use FTS5 query syntax: quotes for an exact phrase, AND, OR and NOT, and pray* for word prefixes.
How it works
Listing. GET /v1/channels/{id}/videos returns 30 videos per page, newest first, at 1 credit per page. The script stops at the first video it already has, so later runs list one page or two. New videos are saved only after the listing finishes, so an interrupted run doesn't leave gaps.
Transcripts. POST /v1/batch takes up to 100 videos at a time, and polling it is free. format: "segments" returns each caption line with its start and end in seconds. The script joins lines into passages of about 45 seconds. Single caption lines are too short to search well; whole videos are too long to point at a moment.
Videos without captions. These come back as 404 and cost nothing. The script marks them no_captions and skips them from then on. YouTube sometimes adds auto-generated captions a while after upload; to retry those, run UPDATE videos SET status = 'new' WHERE status = 'no_captions' in the database and run the script again.
Keeping it current. Run python archive.py on a schedule, weekly with cron for example. Each run lists the newest page and fetches whatever was uploaded since the last one.
What it costs
For a channel with 300 videos, the first run is 10 listing pages and up to 300 transcripts: about 310 credits, less for videos without captions. After that, a weekly run is 1 listing credit plus 1 credit per new video. A $9 pack of 2,000 credits covers the first run and years of updates, and credits don't expire.
Going further
- Language. Without
languages, each transcript comes in the video's own language. For a bilingual channel, add"languages": ["en", "*"]to each task to prefer English where it exists. - Search by meaning. Full-text search finds the words you type. To also find passages about forgiveness that never use the word, embed each passage with an embedding model, and search by similarity alongside FTS5.
- Answer questions. Pass the top passages from
search()to an LLM along with the question, and ask it to answer only from them and cite the links. - More than one channel. Add a
channelcolumn tovideosand loop over a list of handles.
FAQ
Does this work for a small channel with few views?
Yes. YTAPI doesn't need a video to have been requested before, and a small channel's archive is mostly videos nobody has asked for. See new videos and small channels.
What about livestreamed services?
Once a stream has ended and YouTube has processed it, it's a normal video with captions, usually auto-generated. YouTube lists live streams on their own channel tab, so check that the listing includes them. If it doesn't, add their IDs to the videos table yourself.
Can I make the archive public?
The archive holds the transcripts of the channel's videos. For your own channel, or one you have permission from, that's fine. Otherwise keep it for your own use, and link back to the videos rather than republishing the full text.