Next.js · LangChain · OpenAI
Build an AI YouTube Video Summarizer with LangChain, Next.js 15, and YTAPI
A step-by-step tutorial building an AI-powered YouTube summarizer. Learn how to ingest transcripts via REST, chunk content with LangChain, and stream summaries in Next.js 15.
· 3 min read · Platform Team
Most AI YouTube summarizers built on serverless platforms (like Vercel or AWS Amplify) fail in one of two ways:
- They use headless browser scrapers or fragile Python libraries that time out after 10 seconds or get IP-blocked by YouTube.
- They feed messy, unformatted subtitle strings directly into an LLM, wasting thousands of input tokens on redundant timestamps and broken sentences.
In this tutorial, we will build a production-ready YouTube video summarizer using Next.js 15 (App Router), LangChain, OpenAI GPT-4o-mini, and YTAPI.dev.
By requesting transcripts directly in format: "markdown", we get readable paragraphs with timestamp headings in one request, with no scraping or cleanup step before the text goes to the model.
Architectural Overview
Next.js 15 + LangChain Video Summarization Pipeline
Step 1: Project Initialization
Create a new Next.js 15 project with TypeScript and Tailwind CSS:
npx create-next-app@latest yt-summarizer --typescript --tailwind --app --eslint
cd yt-summarizerInstall LangChain, OpenAI SDK, and utility icons:
npm install @langchain/openai @langchain/core lucide-react react-markdownConfigure your environment variables in .env.local:
YT_API_KEY="your_ytapi_key_here"
OPENAI_API_KEY="your_openai_api_key_here"Step 2: Ingesting Transcripts via REST (lib/ytapi.ts)
Instead of parsing raw timed captions, we call the POST /v1/transcripts endpoint with format: "markdown". The response body is the transcript as Markdown text (text/markdown), not JSON: paragraphs under ### [mm:ss] headings, and the video's chapters as ## headings when it has them.
Create lib/ytapi.ts:
export function extractVideoId(urlOrId: string): string | null {
const regExp = /(?:youtube\.com\/(?:[^\/]+\/.+\/|(?:v|e(?:mbed)?)\/|.*[?&]v=)|youtu\.be\/)([^"&?\/\s]{11})/;
const match = urlOrId.match(regExp);
return match ? match[1] : urlOrId.length === 11 ? urlOrId : null;
}
export async function fetchTranscriptMarkdown(videoId: string): Promise<string> {
const apiKey = process.env.YT_API_KEY;
if (!apiKey) {
throw new Error("Missing YT_API_KEY environment variable");
}
const response = await fetch("https://api.ytapi.dev/v1/transcripts", {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
video_id: videoId,
format: "markdown",
}),
// Enable edge caching for repeat queries
next: { revalidate: 86400 },
});
if (!response.ok) {
// Errors are JSON: {"error": {"code": "...", "message": "..."}}.
// A 404 (no captions, private or removed video) is not billed.
const err = await response.json().catch(() => ({}));
throw new Error(err.error?.message ?? `Failed to fetch transcript: ${response.status}`);
}
// format "markdown" returns the Markdown itself, not a JSON envelope.
return response.text();
}Step 3: LangChain Summarization Chain (lib/summarizer.ts)
For short to medium videos (under 12,000 tokens), a single prompt with structured output yields high-density insights. For long lectures or podcasts, we split text into chunks before synthesizing.
Create lib/summarizer.ts:
import { ChatOpenAI } from "@langchain/openai";
import { PromptTemplate } from "@langchain/core/prompts";
import { StringOutputParser } from "@langchain/core/output_parsers";
import { RecursiveCharacterTextSplitter } from "@langchain/textsplitters";
const llm = new ChatOpenAI({
modelName: "gpt-4o-mini",
temperature: 0.3,
streaming: true,
});
const SUMMARY_PROMPT = `
You are an expert executive research assistant. Summarize the following YouTube video transcript into clear, actionable notes.
Structure your response with:
1. **Executive Summary**: 2-3 concise sentences outlining the core thesis.
2. **Key Takeaways**: 4-6 bullet points of specific concepts, data points, or arguments.
3. **Chronological Topic Breakdown**: Major sections with time references if provided in the text.
Transcript:
{transcript}
`;
export async function generateSummaryStream(transcript: string) {
// If transcript is very large, split and combine
if (transcript.length > 35000) {
const splitter = new RecursiveCharacterTextSplitter({
chunkSize: 12000,
chunkOverlap: 1000,
});
const docs = await splitter.createDocuments([transcript]);
// Truncate or map-reduce if needed
transcript = docs.map((d) => d.pageContent).slice(0, 3).join("\n\n---\n\n");
}
const prompt = PromptTemplate.fromTemplate(SUMMARY_PROMPT);
const chain = prompt.pipe(llm).pipe(new StringOutputParser());
return await chain.stream({
transcript,
});
}Step 4: Streaming API Route (app/api/summarize/route.ts)
Create a Server-Sent Events (SSE) stream endpoint to deliver LLM tokens to the frontend with zero perceived delay:
import { NextRequest, NextResponse } from "next/server";
import { extractVideoId, fetchTranscriptMarkdown } from "@/lib/ytapi";
import { generateSummaryStream } from "@/lib/summarizer";
export const runtime = "edge"; // Run on edge for minimum TTFB
export async function POST(req: NextRequest) {
try {
const { url } = await req.json();
if (!url) {
return NextResponse.json({ error: "Missing YouTube URL" }, { status: 400 });
}
const videoId = extractVideoId(url);
if (!videoId) {
return NextResponse.json({ error: "Invalid YouTube URL or ID" }, { status: 400 });
}
// 1. Fetch the transcript as Markdown
const transcript = await fetchTranscriptMarkdown(videoId);
// 2. Stream LangChain response
const stream = await generateSummaryStream(transcript);
const encoder = new TextEncoder();
const readable = new ReadableStream({
async start(controller) {
for await (const chunk of stream) {
controller.enqueue(encoder.encode(chunk));
}
controller.close();
},
});
return new Response(readable, {
headers: {
"Content-Type": "text/plain; charset=utf-8",
"Cache-Control": "no-cache",
},
});
} catch (error: any) {
console.error("Summarization error:", error);
return NextResponse.json(
{ error: error.message || "Failed to summarize video" },
{ status: 500 }
);
}
}Step 5: Frontend Interface (app/page.tsx)
Build a clean UI with real-time text streaming:
"use client";
import { useState } from "react";
import ReactMarkdown from "react-markdown";
import { Loader2, PlayCircle, Sparkles } from "lucide-react";
export default function SummarizerPage() {
const [url, setUrl] = useState("");
const [summary, setSummary] = useState("");
const [loading, setLoading] = useState(false);
const [error, setError] = useState<string | null>(null);
const handleSummarize = async (e: React.FormEvent) => {
e.preventDefault();
setLoading(true);
setError(null);
setSummary("");
try {
const res = await fetch("/api/summarize", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ url }),
});
if (!res.ok) {
const data = await res.json();
throw new Error(data.error || "Failed to generate summary");
}
const reader = res.body?.getReader();
if (!reader) return;
const decoder = new TextDecoder();
let done = false;
while (!done) {
const { value, done: readerDone } = await reader.read();
done = readerDone;
if (value) {
setSummary((prev) => prev + decoder.decode(value));
}
}
} catch (err: any) {
setError(err.message);
} finally {
setLoading(false);
}
};
return (
<main className="min-h-screen bg-slate-950 text-slate-100 p-6 md:p-12 flex flex-col items-center">
<div className="max-w-3xl w-full">
<div className="flex items-center gap-2 mb-2 text-orange-500 font-semibold text-sm">
<Sparkles className="size-4" />
<span>Next.js 15 & LangChain Architecture</span>
</div>
<h1 className="text-3xl font-bold tracking-tight mb-4">
YouTube AI Video Summarizer
</h1>
<p className="text-slate-400 mb-8">
Extract high-density summaries from any YouTube video in seconds using structured markdown transcripts.
</p>
<form onSubmit={handleSummarize} className="flex gap-2 mb-8">
<input
type="text"
placeholder="Paste YouTube video URL (e.g. https://youtu.be/dQw4w9WgXcQ)"
value={url}
onChange={(e) => setUrl(e.target.value)}
required
className="flex-1 px-4 py-3 bg-slate-900 border border-slate-800 rounded-lg text-slate-100 focus:outline-none focus:border-orange-500 text-sm"
/>
<button
type="submit"
disabled={loading}
className="px-6 py-3 bg-orange-600 hover:bg-orange-500 disabled:opacity-50 text-white rounded-lg font-medium text-sm flex items-center gap-2 transition"
>
{loading ? <Loader2 className="size-4 animate-spin" /> : <PlayCircle className="size-4" />}
Summarize
</button>
</form>
{error && (
<div className="p-4 bg-red-950/50 border border-red-800 text-red-200 rounded-lg text-sm mb-6">
{error}
</div>
)}
{summary && (
<div className="bg-slate-900/60 border border-slate-800 rounded-xl p-6 shadow-xl prose prose-invert max-w-none text-slate-200">
<ReactMarkdown>{summary}</ReactMarkdown>
</div>
)}
</div>
</main>
);
}Using LangChain in Python instead
LangChain's Python YoutubeLoader fetches captions with the youtube-transcript-api library, so it breaks the same way once it runs on a server: RequestBlocked or IpBlocked, because YouTube refuses requests from cloud IP ranges. Why that happens on AWS, and on Vercel, Render, and Railway.
You don't need a package to swap it out. A loader is a class with one method:
import os
import requests
from langchain_core.document_loaders import BaseLoader
from langchain_core.documents import Document
class YTAPILoader(BaseLoader):
def __init__(self, video_id: str, languages: str = "en,*"):
self.video_id, self.languages = video_id, languages
def lazy_load(self):
res = requests.post(
"https://api.ytapi.dev/v1/transcripts",
headers={"Authorization": f"Bearer {os.environ['YTAPI_KEY']}"},
json={"video_id": self.video_id, "languages": self.languages.split(","), "format": "markdown"},
timeout=30,
)
res.raise_for_status()
yield Document(page_content=res.text, metadata={"source": f"https://www.youtube.com/watch?v={self.video_id}"})
docs = YTAPILoader("dQw4w9WgXcQ").load()It plugs into text splitters, vector stores and chains like any other loader. Use "format": "text" if you want plain text without timestamp headings.
Production Takeaways
- Format Optimization: Requesting
format: "markdown"cuts token pre-processing latency to zero. Your serverless functions don't spend execution cycles stripping XML or converting JSON arrays into text. - Serverless Execution Windows: Transcript retrieval usually takes well under a second (644 ms at the median for uncached videos in our benchmark), which leaves most of a serverless function's time budget for the model call.
- Edge Caching: Video transcripts rarely mutate. Take advantage of Next.js
fetch(..., { next: { revalidate: 86400 } })to serve subsequent summary requests out of edge cache without consuming additional API credits.
FAQ
Why does LangChain's YoutubeLoader fail on my server but not locally?
It uses youtube-transcript-api, which calls YouTube's caption endpoints from your server's IP. YouTube blocks most datacenter and serverless IP ranges, so the same code that works at home fails in production. A hosted transcript API, like the loader above, moves the fetching off your server.
Which transcript format is best for an LLM?
markdown for summaries and chat: it is readable paragraphs with timestamp headings, with the caption duplicates and line breaks already cleaned up. Use segments if you need start and end times per line for retrieval, and text if you want the words only.
What happens with videos that have no captions?
The API returns a 404 with a code such as captions_disabled, and it is not billed. See what to do when a video has no captions.