Back to blog

Performance · Architecture · Data Engineering

Handling Long YouTube Video Transcripts at Scale: Payloads, Formats, and Memory Bottlenecks

Deep-dive into the architectural mechanics of processing multi-hour YouTube transcripts. Compare XML/SRV1 vs JSON3 payload footprints, streaming JSON parsing, and GC memory stabilization.

· 4 min read · Platform Team

Most developers testing YouTube transcript extraction start with a standard 5-minute tech review or music video. The JSON payload is small (15–45 KB), downloads in 20 milliseconds, and parses without noticeable memory footprint.

Then your ingestion pipeline hits a 3-hour podcast, an all-day developer conference livestream, or an audio-book replay.

Suddenly, your worker processes crash with memory exhaustion:

  • In Node.js: FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory.
  • In Python: Workers killed silently by the Linux kernel OOM-killer (Out of Memory: Kill process).
  • In Go: Heap allocations balloon from 30 MB to several gigabytes under high concurrency.

Understanding how YouTube serializes timed captions, why payload formats differ by 20x, and how to stream large JSON structures is essential for operating high-throughput extraction pipelines.


1. The Video Duration & Payload Scaling Curve

Captions are discrete temporal events. As video duration grows linearly, the number of events, words, and timing objects compounds:

Duration CategoryLength RangeShare of TrafficRaw JSON3 PayloadSentence Segments PayloadMarkdown Payload
ShortUnder 10 minutes~64.9%15 KB – 45 KB6 KB – 18 KB3 KB – 8 KB
Medium10 – 30 minutes~27.1%100 KB – 250 KB35 KB – 80 KB15 KB – 40 KB
Long30 – 60 minutes~5.1%300 KB – 700 KB100 KB – 220 KB50 KB – 110 KB
Extra Long1 – 3+ hours~3.0%800 KB – 2.5+ MB250 KB – 600 KB120 KB – 300 KB

When running 100 concurrent workers ingesting 2-hour podcasts, receiving and unmarshaling 2.5 MB raw JSON3 documents creates 250 MB of raw JSON text transfer and up to 1.5 GB of transient heap objects in memory during deserialization.

Short Video (<10m):   [====] 25 KB
Medium Video (20m):   [================] 150 KB
Extra Long (2.5h):    [========================================================================] 2.2 MB

2. Upstream Caption Formats: SRV1 (XML) vs. JSON3

Under the hood, YouTube’s timedtext infrastructure provides multiple serialization options:

A. Format SRV1 (srv1 - TimedText XML)

The older XML-based protocol. It encapsulates caption snippets into <text> tags:

<?xml version="1.0" encoding="utf-8" ?>
<transcript>
  <text start="12.34" dur="3.45">Good morning everyone, welcome to the talk.</text>
  <text start="15.79" dur="2.10">Today we are discussing distributed systems.</text>
</transcript>
  • Pros: Highly compact. For a 2-hour video, the raw XML transfer is only ~120 KB to 250 KB.
  • Cons: No word-level timing precision; only phrase-level boundaries.

B. Format JSON3 (json3 - Rich Event Tree)

The modern JSON3 protocol used by desktop and mobile clients:

{
  "wireMagic": "pb3",
  "events": [
    {
      "tStartMs": 12340,
      "dDurationMs": 3450,
      "segs": [
        { "utf8": "Good " },
        { "utf8": "morning ", "tOffsetMs": 420 },
        { "utf8": "everyone, ", "tOffsetMs": 850 },
        { "utf8": "welcome ", "tOffsetMs": 1320 }
      ]
    }
  ]
}
  • Pros: Sub-second word alignment via tOffsetMs.
  • Cons: Massive serialization overhead. Each word requires a JSON object with keys and string escapes. For a 3-hour video with 30,000 words, the document exceeds 2 MB.

3. The 80% Efficiency Rule: Request Only What You Need

If your downstream consumer is an LLM (such as GPT-4o or Claude 3.5 Sonnet), requesting format=word_timestamps or raw json3 wastes immense network bandwidth and compute resources.

Comparing Downstream Formats:

Pipeline A (Wasteful):
YTAPI (format=json3) ──> 2.5MB JSON ──> Custom Python Parser ──> 120KB Plain Text ──> LLM

Pipeline B (Optimized):
YTAPI (format=markdown) ──> 140KB Markdown ──> LLM Direct
  • Use format=markdown: The API converts timed events into clean paragraphs with time headers. The payload is much smaller than the raw event format, which cuts network transfer and parsing time.

    The summarizer tutorial uses it end to end, from the API call to a streamed summary.

  • Use format=segments: If you need structured JSON with start and end times for sentence-level RAG indexing, format=segments strips internal word arrays, halving memory usage.

  • Reserve format=word_timestamps strictly for video editors: Only pass word_level=true when you genuinely need millisecond-level word cuts for audio/video synchronization.


4. Memory-Efficient Streaming in Python with ijson

When you must process large word-level JSON files in batch pipelines, avoid json.loads(response.content). That method allocates the full string into RAM and constructs a massive Python dictionary tree simultaneously.

Instead, stream the HTTP response socket directly into an iterative JSON parser like ijson:

pip install ijson httpx
import httpx
import ijson

def stream_large_transcript(video_id: str, api_key: str):
    """
    Stream and process a massive YouTube transcript segment-by-segment 
    without holding the complete JSON payload in memory.
    """
    url = "https://api.ytapi.dev/v1/transcripts"
    params = {
        "video_id": video_id,
        "format": "word_timestamps",
        "word_level": "true"
    }
    headers = {"Authorization": f"Bearer {api_key}"}

    with httpx.Client() as client:
        with client.stream("GET", url, params=params, headers=headers) as response:
            response.raise_for_status()
            
            # ijson.items streams individual segment objects as they arrive over the socket
            segments = ijson.items(response, "segments.item")
            
            total_words = 0
            for segment in segments:
                # Process each segment independently
                text = segment.get("text", "")
                words = segment.get("words", [])
                total_words += len(words)
                
                # Emit or store to database immediately
                if segment.get("start", 0) > 3600:
                    # Example: Process events beyond the 1-hour mark
                    pass

            print(f"Finished streaming video {video_id}. Total words: {total_words}")

Memory Impact:

  • Standard json.loads(): Peaks at 45 MB – 80 MB RAM per 3-hour transcript.
  • Iterative ijson stream: Operates with a flat under 4 MB memory footprint, allowing 100 concurrent workers to run comfortably on a 2GB RAM container.

5. Streaming in Node.js / TypeScript

In Node.js 20+, use web streams with stream/consumers or event emitters to avoid heap spikes:

import { Readable } from "node:stream";

async function processLargeTranscriptStream(videoId: string) {
  const response = await fetch(
    `https://api.ytapi.dev/v1/transcripts?video_id=${videoId}&format=segments`,
    {
      headers: { Authorization: `Bearer ${process.env.YT_API_KEY}` },
    }
  );

  if (!response.ok || !response.body) {
    throw new Error(`Extraction failed: ${response.statusText}`);
  }

  // Convert Web ReadableStream to Node.js Readable
  const nodeStream = Readable.fromWeb(response.body as any);

  let totalBytes = 0;
  nodeStream.on("data", (chunk: Buffer) => {
    totalBytes += chunk.length;
    // Process stream chunks or pipe directly to storage (S3 / disk)
  });

  nodeStream.on("end", () => {
    console.log(`Successfully ingested ${totalBytes} bytes for ${videoId}`);
  });
}

6. Architecture Checklist for Scale

When architecting a pipeline that ingests tens of thousands of YouTube transcripts per day:

  1. Duration Profiling: Treat short (under 10m) and extra-long (over 1h) videos as different workloads in your worker queue. Split long videos into dedicated async worker pools to prevent short jobs from queue head-of-line blocking.
  2. Payload Negotiation: Never default to word_level=true across all jobs. Use format=markdown or format=segments unless downstream algorithms explicitly require word offsets.
  3. Socket Timeouts: Set client-side HTTP timeouts to at least 15–20 seconds for 3-hour videos on cold cache misses.
  4. Cache Verification: Inspect the X-Cache header returned by the API (HIT or MISS). Once a video has been fetched, later requests for it are served from cache and come back faster.

FAQ

Does a long video cost more credits?

No. A transcript costs one credit whatever the video's length, and only successful responses are billed. Long videos cost more in bandwidth and parsing on your side, which is what the format choice above is about.

How do I fetch transcripts for thousands of videos?

Send them as batch jobs of up to 100 tasks each and poll for results, or run your own worker pool against the single-video endpoint. Keep long videos in their own queue so they don't hold up short ones. For a whole channel or playlist, list the video IDs first with the channel videos or playlist endpoints.

Should I request word-level timestamps?

Only if something downstream needs word offsets, such as an editor or karaoke-style captions. Word timings multiply the payload several times over. For search, retrieval and LLM input, segments or markdown carry everything you need.